Skip to content

Commit 4b2ab0d

Browse files
committed
[SPARK-60108][SQL] Scope docs to from_json, fix PERMISSIVE wording, test container values
1 parent 87bfa8c commit 4b2ab0d

2 files changed

Lines changed: 27 additions & 1 deletion

File tree

‎docs/sql-migration-guide.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -28,7 +28,7 @@ license: |
2828

2929
## Upgrading from Spark SQL 4.3 to 4.4
3030

31-
- Since Spark 4.4, when `spark.sql.charVarchar.standardSemantics.enabled` is true, JSON object names used as `MAP<CHAR(n), _>` or `MAP<VARCHAR(n), _>` keys in `from_json` and the JSON datasource are length-checked without padding or trimming. A `CHAR(n)` key must already be exactly `n` characters, and a `VARCHAR(n)` key must already be at most `n` characters. A mismatched key turns the row into a bad record: in `PERMISSIVE` mode it is dropped (the output is `null`), while in `FAILFAST` mode parsing fails with `UNSUPPORTED_JSON_CHAR_VARCHAR_MAP_KEY` (SQLSTATE `0A000`) as the cause of `MALFORMED_RECORD_IN_PARSING`. Duplicate names are kept, exactly as for STRING keys; `spark.sql.mapKeyDedupPolicy` is not applied.
31+
- Since Spark 4.4, when `spark.sql.charVarchar.standardSemantics.enabled` is true, `from_json` length-checks JSON object names used as `MAP<CHAR(n), _>` or `MAP<VARCHAR(n), _>` keys without padding or trimming them. A `CHAR(n)` key must already be exactly `n` characters, and a `VARCHAR(n)` key must already be at most `n` characters; duplicate names are kept, exactly as for STRING keys, and `spark.sql.mapKeyDedupPolicy` is not applied. A key that fails the check makes the enclosing map a bad record: in `PERMISSIVE` mode the map is set to `null` while sibling fields are preserved (the whole `from_json` result is `null` only when the map is the top-level type), and in `FAILFAST` mode parsing fails with `UNSUPPORTED_JSON_CHAR_VARCHAR_MAP_KEY` (SQLSTATE `0A000`) as the cause of `MALFORMED_RECORD_IN_PARSING`. The same parser check runs when reading JSON files with a user-specified CHAR/VARCHAR reader schema. Reads whose CHAR/VARCHAR keys are handled read-side instead (for example catalog tables) are unchanged: there keys are padded, `spark.sql.mapKeyDedupPolicy` is applied, and an over-long key raises `EXCEED_LIMIT_LENGTH`.
3232
- Since Spark 4.4, when `spark.sql.preserveCharVarcharTypeInfo` is true and `spark.sql.charVarchar.standardSemantics.enabled` is false, ORC reads that apply a CHAR/VARCHAR schema over STRING storage return the stored values without ORC truncation, matching Parquet. Previously the ORC reader requested `char(n)`/`varchar(n)` and truncated STRING-stored values to `n`. Read-side length checks (`EXCEED_LIMIT_LENGTH`) apply only when `spark.sql.charVarchar.standardSemantics.enabled` is true.
3333
- Since Spark 4.4, the options maps passed to `from_csv`, `to_csv`, `schema_of_csv`, `from_json`, `to_json`, `schema_of_json`, `from_xml`, `to_xml`, and `schema_of_xml` must be foldable after replacing `RuntimeReplaceable` expressions. Previously, Spark evaluated non-foldable options during analysis, which allowed some constant expressions but could fail with an internal error or incorrectly evaluate row-dependent expressions. To allow deterministic and row-independent non-foldable options, set `spark.sql.legacy.allowNonFoldableOptions` to `true`. Row-dependent, unevaluable, and nondeterministic options are always rejected.
3434
- Since Spark 4.4, when an already-analyzed Data Source V2 query is refreshed after a compatible schema change, connectors can return more data columns from the current table schema in `Scan.readSchema()` than requested by `SupportsPushDownRequiredColumns.pruneColumns`. Previously, this partial pruning could fail planning because the scan reported columns absent from the analyzed relation output.

‎sql/core/src/test/scala/org/apache/spark/sql/CharVarcharTestSuite.scala‎

Lines changed: 26 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -3235,6 +3235,32 @@ class BasicCharVarcharTestSuite extends SharedSparkSession {
32353235
}
32363236
}
32373237

3238+
// A rejected key whose value is an object or array exercises skipChildren(): the entire
3239+
// nested value is consumed so the map is nulled and the sibling `tail` field survives, and
3240+
// in FAILFAST the key is still surfaced. This holds in both partial-results modes.
3241+
Seq(true, false).foreach { partial =>
3242+
withSQLConf(SQLConf.JSON_ENABLE_PARTIAL_RESULTS.key -> partial.toString) {
3243+
checkAnswer(
3244+
sql("""SELECT from_json('{"m":{"a":{"x":1}},"tail":2}',
3245+
| 'm MAP<CHAR(3), STRUCT<x: INT>>, tail INT')""".stripMargin),
3246+
Row(Row(null, 2)))
3247+
checkAnswer(
3248+
sql("""SELECT from_json('{"m":{"a":[1,2,3]},"tail":2}',
3249+
| 'm MAP<CHAR(3), ARRAY<INT>>, tail INT')""".stripMargin),
3250+
Row(Row(null, 2)))
3251+
assertUnsupportedJsonMapKey(
3252+
"""SELECT from_json('{"a":{"x":1}}',
3253+
| 'MAP<CHAR(3), STRUCT<x: INT>>', map('mode', 'FAILFAST'))""".stripMargin,
3254+
key = "a",
3255+
dataTypeSql = "CHAR(3)")
3256+
assertUnsupportedJsonMapKey(
3257+
"""SELECT from_json('{"a":[1,2,3]}',
3258+
| 'MAP<CHAR(3), ARRAY<INT>>', map('mode', 'FAILFAST'))""".stripMargin,
3259+
key = "a",
3260+
dataTypeSql = "CHAR(3)")
3261+
}
3262+
}
3263+
32383264
// multiLine + streaming top-level array: a bad key in one element must not add spurious
32393265
// rows. Two elements in, two rows out, with the bad element reported in _corrupt_record.
32403266
withTempPath { file =>

0 commit comments

Comments
 (0)