Skip to content

Commit 8e1c55b

Browse files
committed
[SPARK-60103][SQL] Keep XML MAP struct fields on convertField dispatch
Drop the direct convertObject MapType arm so empty, attribute-only, and text-only MAP<STRING> elements stay null or malformed. Restrict convertMap to StringType-family keys so format("xml") MAP<INT> does not ClassCast.
1 parent 02b478e commit 8e1c55b

3 files changed

Lines changed: 65 additions & 17 deletions

File tree

‎docs/sql-migration-guide.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -29,7 +29,7 @@ license: |
2929
## Upgrading from Spark SQL 4.3 to 4.4
3030

3131
- Since Spark 4.4, when `spark.sql.preserveCharVarcharTypeInfo` is true and `spark.sql.charVarchar.standardSemantics.enabled` is false, ORC reads that apply a CHAR/VARCHAR schema over STRING storage return the stored values without ORC truncation, matching Parquet. Previously the ORC reader requested `char(n)`/`varchar(n)` and truncated STRING-stored values to `n`. Read-side length checks (`EXCEED_LIMIT_LENGTH`) apply only when `spark.sql.charVarchar.standardSemantics.enabled` is true.
32-
- Since Spark 4.4, when `spark.sql.charVarchar.standardSemantics.enabled` is true, XML element and attribute names used as `MAP<CHAR(n), _>` or `MAP<VARCHAR(n), _>` keys in `from_xml` and the XML datasource are length-checked without padding or trimming. A `CHAR(n)` key must already be exactly `n` characters, and a `VARCHAR(n)` key must already be at most `n` characters. Mismatched keys fail with `UNSUPPORTED_XML_CHAR_VARCHAR_MAP_KEY` (SQLSTATE `0A000`) and follow the XML parse mode (`PERMISSIVE` or `FAILFAST`). Exact repeated names last-win; `spark.sql.mapKeyDedupPolicy` is not applied. XML CHAR/VARCHAR values still use write-side pad and `EXCEED_LIMIT_LENGTH`. With the flag off, CHAR keys are still padded and VARCHAR overflow still uses `EXCEED_LIMIT_LENGTH`.
32+
- Since Spark 4.4, when `spark.sql.charVarchar.standardSemantics.enabled` is true, XML names used as `MAP<CHAR(n), _>` or `MAP<VARCHAR(n), _>` keys in `from_xml` and the XML datasource are length-checked without padding or trimming. The checked names are element names, `attributePrefix` plus attribute names (default prefix `_`), and the `valueTag` (default `_VALUE`) when mixed text is present. A `CHAR(n)` key must already be exactly `n` characters, and a `VARCHAR(n)` key must already be at most `n` characters. Mismatched keys fail with `UNSUPPORTED_XML_CHAR_VARCHAR_MAP_KEY` (SQLSTATE `0A000`) and follow the XML parse mode (`PERMISSIVE` or `FAILFAST`). Exact repeated names last-win; `spark.sql.mapKeyDedupPolicy` is not applied. XML CHAR/VARCHAR values use the same pad and `EXCEED_LIMIT_LENGTH` checks as other parsed text. Empty, attribute-only, and text-only map elements stay SQL NULL or a malformed record, matching 4.3; they do not become an empty map or a `valueTag` entry. In Spark 4.3, `MAP<CHAR(n), _>` and `MAP<VARCHAR(n), _>` XML maps were malformed records because only unbounded STRING keys were accepted. Padding of CHAR keys and VARCHAR `EXCEED_LIMIT_LENGTH` apply when the standard-semantics flag is false and `spark.sql.preserveCharVarcharTypeInfo` is true, so first-class CHAR/VARCHAR types still reach the parser.
3333
- Since Spark 4.4, the options maps passed to `from_csv`, `to_csv`, `schema_of_csv`, `from_json`, `to_json`, `schema_of_json`, `from_xml`, `to_xml`, and `schema_of_xml` must be foldable after replacing `RuntimeReplaceable` expressions. Previously, Spark evaluated non-foldable options during analysis, which allowed some constant expressions but could fail with an internal error or incorrectly evaluate row-dependent expressions. To allow deterministic and row-independent non-foldable options, set `spark.sql.legacy.allowNonFoldableOptions` to `true`. Row-dependent, unevaluable, and nondeterministic options are always rejected.
3434
- Since Spark 4.4, when an already-analyzed Data Source V2 query is refreshed after a compatible schema change, connectors can return more data columns from the current table schema in `Scan.readSchema()` than requested by `SupportsPushDownRequiredColumns.pruneColumns`. Previously, this partial pruning could fail planning because the scan reported columns absent from the analyzed relation output.
3535
- Since Spark 4.4, when an already-analyzed Data Source V2 query is refreshed, if rebinding an expanded struct used as a map key causes distinct current keys to collide under the analyzed schema, the query fails with `DUPLICATED_MAP_KEY` by default instead of returning duplicate keys. With `spark.sql.mapKeyDedupPolicy=LAST_WIN`, the last value is kept.

‎sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/xml/StaxXmlParser.scala‎

Lines changed: 13 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -385,7 +385,10 @@ class StaxXmlParser(
385385
startElementName: String,
386386
attributes: Array[Attribute]): Any = dt match {
387387
case st: StructType => convertObject(parser, st)
388-
case MapType(kt, vt, _) => convertMap(parser, vt, attributes, kt)
388+
// CHAR/VARCHAR extend StringType. Non-string keys (for example MAP<INT, _> through
389+
// format("xml").load()) must not reach convertMap: applyTextParseSemantics would keep
390+
// UTF8String keys and fail later with ClassCastException.
391+
case MapType(kt: StringType, vt, _) => convertMap(parser, vt, attributes, kt)
389392
case ArrayType(st, _) => convertField(parser, st, startElementName)
390393
case VariantType =>
391394
StaxXmlParser.convertVariant(parser, attributes, options)
@@ -447,14 +450,18 @@ class StaxXmlParser(
447450
/**
448451
* Parse an object as a Map.
449452
*
450-
* XML names used as CHAR/VARCHAR keys are length-checked without rewriting:
451-
* CHAR keys must already be exactly n characters, and VARCHAR keys must already be
452-
* at most n characters. Padding, trimming, and mapKeyDedupPolicy are not applied.
453+
* When `spark.sql.charVarchar.standardSemantics.enabled` is true, XML names used as
454+
* CHAR/VARCHAR keys are length-checked without rewriting: CHAR keys must already be
455+
* exactly n characters, and VARCHAR keys must already be at most n characters.
456+
* Padding, trimming, and mapKeyDedupPolicy are not applied. With the flag off,
457+
* CHAR keys are padded and VARCHAR overflow uses EXCEED_LIMIT_LENGTH when
458+
* first-class CHAR/VARCHAR types reach the parser.
453459
* Repeated names last-win on binary equality (collation is not consulted), matching
454460
* ordinary MAP<STRING, ...> XML maps.
455461
*
456-
* Example: `from_xml('<ROW><m><ab>1</ab></m></ROW>', 'm MAP<CHAR(2), INT>')`
457-
* keeps key `ab`; `MAP<CHAR(4), INT>` raises `UNSUPPORTED_XML_CHAR_VARCHAR_MAP_KEY`.
462+
* Example (flag on): `from_xml('<ROW><m><ab>1</ab></m></ROW>',
463+
* 'm MAP<CHAR(2), INT>')` keeps key `ab`; `MAP<CHAR(4), INT>` raises
464+
* `UNSUPPORTED_XML_CHAR_VARCHAR_MAP_KEY`.
458465
*
459466
* This method owns element bounding so a key-check failure still consumes the current
460467
* map element (ARRAY<MAP<...>> and nested maps included).
@@ -647,9 +654,6 @@ class StaxXmlParser(
647654
case st: StructType =>
648655
row(index) = convertNestedStruct(parser, st, field, attributes)
649656

650-
case mt: MapType =>
651-
row(index) = convertMap(parser, mt.valueType, attributes, mt.keyType)
652-
653657
case ArrayType(dt: DataType, _) =>
654658
val values = Option(row(index))
655659
.map(_.asInstanceOf[ArrayBuffer[Any]])

‎sql/core/src/test/scala/org/apache/spark/sql/CharVarcharTestSuite.scala‎

Lines changed: 51 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -2978,13 +2978,14 @@ class BasicCharVarcharTestSuite extends SharedSparkSession {
29782978
key = "a",
29792979
dataTypeSql = "CHAR(2)")
29802980

2981-
// Default valueTag is "_VALUE" (6 characters), so CHAR(2) rejects it.
2981+
// Default valueTag is "_VALUE" (6 characters). Text-only elements are malformed
2982+
// records (convertField), so use mixed content to reach convertMap.
29822983
checkAnswer(
2983-
sql("SELECT from_xml('<ROW><m>xy</m></ROW>', 'm MAP<CHAR(2), STRING>')"),
2984+
sql("SELECT from_xml('<ROW><m><ab>1</ab>xy</m></ROW>', 'm MAP<CHAR(2), STRING>')"),
29842985
Row(Row(null)))
29852986
assertUnsupportedXmlMapKey(
29862987
"""SELECT from_xml(
2987-
| '<ROW><m>xy</m></ROW>',
2988+
| '<ROW><m><ab>1</ab>xy</m></ROW>',
29882989
| 'm MAP<CHAR(2), STRING>',
29892990
| map('mode', 'FAILFAST'))""".stripMargin,
29902991
key = "_VALUE",
@@ -3049,20 +3050,38 @@ class BasicCharVarcharTestSuite extends SharedSparkSession {
30493050
dataTypeSql = "VARCHAR(2)")
30503051

30513052
// Default attributePrefix "_" is part of the key: local name ab is key _ab.
3053+
// Attribute-only elements are SQL NULL (convertField EndElement), so include a child.
30523054
checkAnswer(
3053-
sql("SELECT from_xml('<ROW><m ab=\"1\"></m></ROW>', 'm MAP<CHAR(2), INT>')"),
3055+
sql("SELECT from_xml('<ROW><m ab=\"1\"><xy>2</xy></m></ROW>', 'm MAP<CHAR(2), INT>')"),
30543056
Row(Row(null)))
30553057
checkAnswer(
3056-
sql("SELECT from_xml('<ROW><m ab=\"1\"></m></ROW>', 'm MAP<CHAR(3), INT>')"),
3057-
Row(Row(Map("_ab" -> 1))))
3058+
sql("SELECT from_xml('<ROW><m ab=\"1\"><xyz>2</xyz></m></ROW>', 'm MAP<CHAR(3), INT>')"),
3059+
Row(Row(Map("_ab" -> 1, "xyz" -> 2))))
30583060
assertUnsupportedXmlMapKey(
30593061
"""SELECT from_xml(
3060-
| '<ROW><m ab="1"></m></ROW>',
3062+
| '<ROW><m ab="1"><xy>2</xy></m></ROW>',
30613063
| 'm MAP<CHAR(2), INT>',
30623064
| map('mode', 'FAILFAST'))""".stripMargin,
30633065
key = "_ab",
30643066
dataTypeSql = "CHAR(2)")
30653067

3068+
// MAP<STRING> empty, attribute-only, and text-only elements stay SQL NULL, matching
3069+
// convertField on master. ARRAY<MAP> empty elements are null entries, not {}.
3070+
checkAnswer(
3071+
sql("SELECT from_xml('<ROW><m></m></ROW>', 'm MAP<STRING, INT>')"),
3072+
Row(Row(null)))
3073+
checkAnswer(
3074+
sql("SELECT from_xml('<ROW><m a=\"1\"></m></ROW>', 'm MAP<STRING, INT>')"),
3075+
Row(Row(null)))
3076+
checkAnswer(
3077+
sql("SELECT from_xml('<ROW><m>xy</m></ROW>', 'm MAP<STRING, STRING>')"),
3078+
Row(Row(null)))
3079+
checkAnswer(
3080+
sql("""SELECT from_xml(
3081+
| '<ROW><m></m><tail>9</tail></ROW>',
3082+
| 'm ARRAY<MAP<STRING, INT>>, tail INT')""".stripMargin),
3083+
Row(Row(Seq(null), 9)))
3084+
30663085
withTempPath { path =>
30673086
Seq(
30683087
"""<ROWS><ROW><m><ab>1</ab></m></ROW><ROW><m><a>1</a></m></ROW></ROWS>"""
@@ -3099,6 +3118,31 @@ class BasicCharVarcharTestSuite extends SharedSparkSession {
30993118
checkAnswer(arrayPermissive, Seq(Row(null), Row(Seq(Map("ab" -> 2)))))
31003119
}
31013120

3121+
// format("xml").load() skips INVALID_XML_MAP_KEY_TYPE. Non-string keys must stay a
3122+
// malformed record (MatchError), not a mistyped map that fails with ClassCastException.
3123+
withTempPath { path =>
3124+
Seq("""<ROWS><ROW><m><a>x</a></m></ROW></ROWS>""")
3125+
.toDS().write.text(path.getCanonicalPath)
3126+
val intKeySchema = "m MAP<INT, STRING>"
3127+
val permissiveInt = spark.read
3128+
.format("xml")
3129+
.option("rowTag", "ROW")
3130+
.schema(intKeySchema)
3131+
.load(path.getCanonicalPath)
3132+
checkAnswer(permissiveInt, Seq(Row(null)))
3133+
val failFastInt = spark.read
3134+
.format("xml")
3135+
.option("rowTag", "ROW")
3136+
.option("mode", "FAILFAST")
3137+
.schema(intKeySchema)
3138+
.load(path.getCanonicalPath)
3139+
val error = intercept[Exception] { failFastInt.collect() }
3140+
val causes = Iterator.iterate[Throwable](error)(_.getCause).takeWhile(_ != null)
3141+
assert(
3142+
!causes.exists(_.isInstanceOf[ClassCastException]),
3143+
s"MAP<INT, STRING> XML read must not ClassCastException: $error")
3144+
}
3145+
31023146
withSQLConf(
31033147
SQLConf.MAP_KEY_DEDUP_POLICY.key -> SQLConf.MapKeyDedupPolicy.EXCEPTION.toString) {
31043148
checkAnswer(

0 commit comments

Comments
 (0)