Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions common/utils/src/main/resources/error/error-conditions.json
Original file line number Diff line number Diff line change
Expand Up @@ -9605,6 +9605,13 @@
],
"sqlState" : "0A000"
},
"UNSUPPORTED_JSON_CHAR_VARCHAR_MAP_KEY" : {
"message" : [
"The JSON object name <key> is not a valid key for the map key type <dataType>.",
"A CHAR key must match the declared length exactly and a VARCHAR key must not exceed it; keys are never padded or trimmed. Use a STRING map key to accept any name."
],
"sqlState" : "0A000"
},
"UNSUPPORTED_MERGE_CONDITION" : {
"message" : [
"MERGE operation contains unsupported <condName> condition."
Expand Down
1 change: 1 addition & 0 deletions docs/sql-migration-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ license: |

## Upgrading from Spark SQL 4.3 to 4.4

- Since Spark 4.4, when `spark.sql.charVarchar.standardSemantics.enabled` is true, `from_json` length-checks JSON object names used as `MAP<CHAR(n), _>` or `MAP<VARCHAR(n), _>` keys without padding or trimming them. A `CHAR(n)` key must already be exactly `n` characters, and a `VARCHAR(n)` key must already be at most `n` characters; duplicate names are kept, exactly as for STRING keys, and `spark.sql.mapKeyDedupPolicy` is not applied. A key that fails the check makes the enclosing map a bad record: in `PERMISSIVE` mode the map is set to `null` while sibling fields are preserved (the whole `from_json` result is `null` only when the map is the top-level type), and in `FAILFAST` mode parsing fails with `UNSUPPORTED_JSON_CHAR_VARCHAR_MAP_KEY` (SQLSTATE `0A000`) as the cause of `MALFORMED_RECORD_IN_PARSING`. The same parser check runs when reading JSON files with a user-specified CHAR/VARCHAR reader schema. Reads whose CHAR/VARCHAR keys are handled read-side instead (for example catalog tables) are unchanged: there keys are padded, `spark.sql.mapKeyDedupPolicy` is applied, and an over-long key raises `EXCEED_LIMIT_LENGTH`.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking (P2): The last sentence here says that reads whose CHAR/VARCHAR keys are handled read-side, "for example catalog tables", are unchanged: keys are padded, spark.sql.mapKeyDedupPolicy is applied, and long keys raise EXCEED_LIMIT_LENGTH. That is not what happens under standard semantics. The sentence comes from my earlier comment, which had this wrong.

With spark.sql.charVarchar.standardSemantics.enabled=true, charVarcharFirstClassTypes is true, so replaceCharVarcharWithString keeps CharType. For a catalog JSON table, SessionCatalog.getTableMetadata and LogicalRelation.apply keep MAP<CHAR(3), INT>, DataSourceStrategy.readDataSourceTable passes it as the reader schema, and JsonFileFormat builds JacksonParser over it. makeMapKeyChecker then installs the new check:

CREATE TABLE t (m MAP<CHAR(3), INT>) USING json;
-- data rows: {"m":{"a":1}} and {"m":{"abcd":1}}
SELECT m FROM t;
-- PERMISSIVE: m is NULL for both rows
-- FAILFAST: UNSUPPORTED_JSON_CHAR_VARCHAR_MAP_KEY

Neither row gets the documented 'a ' or EXCEED_LIMIT_LENGTH. Keys that pass the parser are then rebuilt by the read-side MapFromArrays, for catalog tables and user-specified reader schemas alike. So {"m":{"abc":1,"abc":2}} still raises DUPLICATED_MAP_KEY under the default policy, unlike from_json. Could this sentence, and the matching PR-description paragraph, say that JSON scans with CHAR/VARCHAR map keys run the same parser check and still apply mapKeyDedupPolicy? If catalog reads should skip the parser check instead, that would need a code change.

See Shared repair plan 1 in the review body.

- Since Spark 4.4, when `spark.sql.preserveCharVarcharTypeInfo` is true and `spark.sql.charVarchar.standardSemantics.enabled` is false, ORC reads that apply a CHAR/VARCHAR schema over STRING storage return the stored values without ORC truncation, matching Parquet. Previously the ORC reader requested `char(n)`/`varchar(n)` and truncated STRING-stored values to `n`. Read-side length checks (`EXCEED_LIMIT_LENGTH`) apply only when `spark.sql.charVarchar.standardSemantics.enabled` is true.
- Since Spark 4.4, the options maps passed to `from_csv`, `to_csv`, `schema_of_csv`, `from_json`, `to_json`, `schema_of_json`, `from_xml`, `to_xml`, and `schema_of_xml` must be foldable after replacing `RuntimeReplaceable` expressions. Previously, Spark evaluated non-foldable options during analysis, which allowed some constant expressions but could fail with an internal error or incorrectly evaluate row-dependent expressions. To allow deterministic and row-independent non-foldable options, set `spark.sql.legacy.allowNonFoldableOptions` to `true`. Row-dependent, unevaluable, and nondeterministic options are always rejected.
- Since Spark 4.4, when an already-analyzed Data Source V2 query is refreshed after a compatible schema change, connectors can return more data columns from the current table schema in `Scan.readSchema()` than requested by `SupportsPushDownRequiredColumns.pruneColumns`. Previously, this partial pruning could fail planning because the scan reported columns absent from the analyzed relation output.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -119,7 +119,8 @@ case class JsonToStructsEvaluator(
nullableSchema: DataType,
nameOfCorruptRecord: String,
timeZoneId: Option[String],
variantAllowDuplicateKeys: Boolean) {
variantAllowDuplicateKeys: Boolean,
charVarcharStandardSemantics: Boolean) {

// This converts parsed rows to the desired output by the given schema.
@transient
Expand Down Expand Up @@ -149,7 +150,8 @@ case class JsonToStructsEvaluator(
(StructType(Array(StructField("value", other))), other)
}

val rawParser = new JacksonParser(actualSchema, parsedOptions, allowArrayAsStructs = false)
val rawParser = new JacksonParser(actualSchema, parsedOptions, allowArrayAsStructs = false,
charVarcharStandardSemanticsOverride = Some(charVarcharStandardSemantics))
val createParser = CreateJacksonParser.utf8String _

new FailureSafeParser[UTF8String](
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1873,7 +1873,14 @@ case class JsonToStructs(
options: Map[String, String],
child: Expression,
timeZoneId: Option[String] = None,
variantAllowDuplicateKeys: Boolean = SQLConf.get.getConf(SQLConf.VARIANT_ALLOW_DUPLICATE_KEYS))
variantAllowDuplicateKeys: Boolean = SQLConf.get.getConf(SQLConf.VARIANT_ALLOW_DUPLICATE_KEYS),
// `spark.sql.charVarchar.standardSemantics.enabled` has PERSISTED binding, so it is captured
// here at analysis time (as the default arg, like `variantAllowDuplicateKeys`) and then
// threaded all the way into the parser via `charVarcharStandardSemanticsOverride`. This keeps
// a view's CHAR/VARCHAR map-key semantics tied to its creation-time flag rather than the
// caller's session setting. (`variantAllowDuplicateKeys` is only captured, not threaded: the
// parser still reads that one live from `SQLConf.get`.)
charVarcharStandardSemantics: Boolean = SQLConf.get.charVarcharStandardSemantics)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit (P3): This default is evaluated wherever JsonToStructs is constructed, and one caller constructs it at execution. BaseScriptTransformationExec.outputFieldWriters is a lazy val. Its non-standard complex-type branch (case _: ArrayType | _: MapType | _: StructType) builds JsonToStructs(attr.dataType, ioschema.outputSerdeProps.toMap, Literal(null), Some(conf.sessionLocalTimeZone)) without this argument, so it picks up SQLConf.get at execution. That file says the mode is "Bound at parse into the I/O schema ... Do not re-read SQLConf".

Take a no-SerDe TRANSFORM ... AS (m MAP<CHAR(3), INT>) parsed with standard semantics off and spark.sql.preserveCharVarcharTypeInfo on. That happens in a persisted view, or in a DataFrame whose conf changes before the action. If it then executes with the flag on, it gets the CHAR(3) key check, and a script row {"a":1} becomes a null map, although the PR description says TRANSFORM is unchanged. Passing charVarcharStandardSemantics = standardCharVarcharSemantics at that construction would keep TRANSFORM on the mode it bound at parse.

Verification:

  • Regression: A no-SerDe TRANSFORM with MAP<CHAR(3), INT> output, parsed with standard semantics off and preserveCharVarcharTypeInfo on, then executed with standard semantics on, keeps an off-width script output key instead of nulling the map.
  • Compatibility: TRANSFORM output under standard semantics (the standard branch) is unchanged.

extends UnaryExpression
with TimeZoneAwareExpression
with CodegenFallback
Expand Down Expand Up @@ -1935,7 +1942,8 @@ case class JsonToStructs(

@transient
private lazy val evaluator = new JsonToStructsEvaluator(
options, nullableSchema, nameOfCorruptRecord, timeZoneId, variantAllowDuplicateKeys)
options, nullableSchema, nameOfCorruptRecord, timeZoneId, variantAllowDuplicateKeys,
charVarcharStandardSemantics)
override def stateful: Boolean = true

override def nullSafeEval(json: Any): Any = {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -51,7 +51,13 @@ class JacksonParser(
schema: DataType,
val options: JSONOptions,
allowArrayAsStructs: Boolean,
filters: Seq[Filter] = Seq.empty) extends Logging {
filters: Seq[Filter] = Seq.empty,
// `spark.sql.charVarchar.standardSemantics.enabled` has PERSISTED binding, so for `from_json`
// it must be captured when the expression is analyzed (see `JsonToStructs`) rather than read
// live here, otherwise a view created under the flag would skip the CHAR/VARCHAR map-key check
// when queried from a session with the flag off. `None` falls back to the live conf, which is
// correct for the file-based JSON data source where there is no view to capture.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit (P3): "there is no view to capture" does not hold for file scans. A persisted view can read a catalog JSON table, and under standard semantics that table keeps MAP<CHAR(n), _> all the way to the scan. The view body (table resolution plus ApplyCharTypePadding) is analyzed under the view's flag. But JsonFileFormat and JsonPartitionReaderFactory build this parser with None, so the key check reads the querying session's flag at execution:

-- created with the flag on
CREATE TABLE t (m MAP<CHAR(3), INT>) USING json;  -- data row: {"m":{"a":1}}
CREATE VIEW v AS SELECT m FROM t;
SELECT * FROM v;
-- queried with the flag on: m is NULL
-- queried with the flag off: m = {'a  ' -> 1}

In the flag-off session the parser installs no checker and accepts 'a', and the view's read-side rewrite pads it. The same persisted view returns different data depending on the caller. Other JSON scan options already follow the session, so documenting this may be enough. Could this comment, and the guide's datasource sentence, say that the file-scan check uses the executing session's flag, rather than that no view can be involved?

Verification:

  • Behavior: A persisted view over a JSON table with a MAP<CHAR(3), INT> column, created under standard semantics, produces the documented results for an off-width key when queried from flag-on and flag-off sessions.
  • Inspection: The JacksonParser comment and the migration-guide sentence describe the same session-bound file-scan behavior the test asserts.

charVarcharStandardSemanticsOverride: Option[Boolean] = None) extends Logging {

import JacksonUtils._
import com.fasterxml.jackson.core.JsonToken._
Expand All @@ -60,6 +66,13 @@ class JacksonParser(
// to a value in a field for `InternalRow`.
private type ValueConverter = JsonParser => AnyRef

// CHAR/VARCHAR JSON map keys are length-checked without pad or trim under this flag. Lazy so
// its value is independent of field declaration order: `makeMapKeyChecker` reads it while the
// converters below are built, and an eager val read before its own initializer would silently
// see `false` and disable the check.
private lazy val charVarcharStandardSemantics =
charVarcharStandardSemanticsOverride.getOrElse(SQLConf.get.charVarcharStandardSemantics)

// `ValueConverter`s for the root schema for all fields in the schema
private val rootConverter = makeRootConverter(schema)

Expand Down Expand Up @@ -191,8 +204,9 @@ class JacksonParser(

private def makeMapRootConverter(mt: MapType): JsonParser => Iterable[InternalRow] = {
val fieldConverter = makeConverter(mt.valueType)
val keyChecker = makeMapKeyChecker(mt.keyType)
(parser: JsonParser) => parseJsonToken[Iterable[InternalRow]](parser, mt) {
case START_OBJECT => Some(InternalRow(convertMap(parser, fieldConverter)))
case START_OBJECT => Some(InternalRow(convertMap(parser, fieldConverter, keyChecker)))
}
}

Expand Down Expand Up @@ -484,8 +498,9 @@ class JacksonParser(

case mt: MapType =>
val valueConverter = makeConverter(mt.valueType)
val keyChecker = makeMapKeyChecker(mt.keyType)
(parser: JsonParser) => parseJsonToken[MapData](parser, dataType) {
case START_OBJECT => convertMap(parser, valueConverter)
case START_OBJECT => convertMap(parser, valueConverter, keyChecker)
}

case udt: UserDefinedType[_] =>
Expand Down Expand Up @@ -618,38 +633,65 @@ class JacksonParser(
}

/**
* Parse an object as a Map, preserving all fields.
* Parse an object as a Map.
*
* JSON object names used as CHAR/VARCHAR keys are length-checked without rewriting (see
* `makeMapKeyChecker`): CHAR keys must already be exactly n characters, and VARCHAR keys
* must already be at most n characters. Padding, trimming, and `mapKeyDedupPolicy` are not
* applied, and duplicate names are kept, exactly as for STRING keys. This differs from XML,
* which pads keys and then applies `mapKeyDedupPolicy`.
*
* A rejected key is treated like a failed value, regardless of `enablePartialResults`: the
* cause is recorded and the parser steps past the value so the loop still consumes the map's
* END_OBJECT, leaving the key unpaired. The check must not surface its result while the parser
* is still on the FIELD_NAME, or an enclosing struct would misread the map's remaining entries
* as its own sibling fields (SPARK-60108). STRING keys have no checker and take the fast path.
*/
private def convertMap(
parser: JsonParser,
fieldConverter: ValueConverter): MapData = {
fieldConverter: ValueConverter,
keyChecker: Option[UTF8String => Option[Throwable]]): MapData = {
val keys = ArrayBuffer.empty[UTF8String]
val values = ArrayBuffer.empty[Any]
var badRecordException: Option[Throwable] = None

while (nextUntil(parser, JsonToken.END_OBJECT)) {
keys += UTF8String.fromString(parser.currentName)
try {
values += fieldConverter.apply(parser)
} catch {
case err: PartialValueException if enablePartialResults =>
badRecordException = badRecordException.orElse(Some(err.cause))
values += err.partialResult
case NonFatal(e) if enablePartialResults =>
val key = UTF8String.fromString(parser.currentName)
keys += key
// Avoid allocating a closure per entry on the common no-checker path.
val keyRejection = keyChecker match {
case Some(check) => check(key)
case None => None
}
keyRejection match {
case Some(e) =>
badRecordException = badRecordException.orElse(Some(e))
parser.nextToken() // step from FIELD_NAME onto the value
parser.skipChildren()
case None =>
try {
values += fieldConverter.apply(parser)
} catch {
case err: PartialValueException if enablePartialResults =>
badRecordException = badRecordException.orElse(Some(err.cause))
values += err.partialResult
case NonFatal(e) if enablePartialResults =>
badRecordException = badRecordException.orElse(Some(e))
parser.skipChildren()
}
}
}

// Value conversion can fail after the key is recorded. Do not build MapData from
// unpaired buffers: ArrayBasedMapData would throw a cardinality error and hide
// the original conversion failure (for example EXCEED_LIMIT_LENGTH).
// A rejected key or a failed value leaves the key recorded without a value. Rethrow the
// real cause instead of building an unbalanced ArrayBasedMapData, whose cardinality
// `require` would otherwise mask it. The row then becomes a bad record (null in PERMISSIVE,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit (P3): This says the row then becomes a bad record, null in PERMISSIVE mode. In fact the rethrow leaves containment to the enclosing converter. When the map is a struct field, convertObject absorbs it and nulls only that field, which is what the new test expects: Row(Row(null, 1)) for {"m":{"a":1},"tail":1}. An enclosing array or map, or the root, propagates it further. "(a trimmed value)" on the next line also has no counterpart in the code. The balanced partial results that reach PartialMapDataResultException come from nested struct, array, or map values throwing PartialValueException. CHAR/VARCHAR value handling returns the padded or trimmed string without throwing. Maybe something like: "Rethrow the real cause and let the enclosing converter decide containment: a struct nulls just this field, while an array, a map, or the root propagates it. Balanced partial results from nested struct, array, or map values still flow through PartialMapDataResultException below."

Verification:

  • Inspection: The comment matches convertObject/convertArray handling and the PartialValueException producers, consistent with the sibling-preservation test expecting Row(Row(null, 1)).

// surfaced in FAILFAST). Balanced partial results (a trimmed value) still flow through
// PartialMapDataResultException below.
if (keys.length != values.length) {
throw badRecordException.get
}

// Preserve every parsed JSON key/value pair, including exact duplicate names.
// ArrayBasedMapData is used directly to retain this historical behavior.
// The JSON map keeps every parsed pair, including exact duplicate names.
val mapData = ArrayBasedMapData(keys.toArray, values.toArray)

if (badRecordException.isEmpty) {
Expand All @@ -659,6 +701,44 @@ class JacksonParser(
}
}

/**
* Builds the length check applied to this map's CHAR/VARCHAR keys, or `None` when the keys
* need no check (STRING keys, or standard semantics are off). Built once per map converter so
* the per-entry loop in `convertMap` does not re-inspect the key type. The check never pads or
* trims; it returns the rejection cause (with the actual key type, collation included, so the
* message is accurate) rather than throwing, so `convertMap` keeps full control of the parser
* position when a key is rejected.
*/
private def makeMapKeyChecker(keyType: DataType): Option[UTF8String => Option[Throwable]] = {
if (!charVarcharStandardSemantics) {
None
} else {
keyType match {
case c: CharType => Some(checkCharJsonMapKey(_, c))
case v: VarcharType => Some(checkVarcharJsonMapKey(_, v))
case _ => None
}
}
}

// A CHAR(n) JSON object name is never padded: it must already be exactly n characters.
private def checkCharJsonMapKey(key: UTF8String, keyType: CharType): Option[Throwable] = {
if (key.numChars() != keyType.length) {
Some(QueryExecutionErrors.unsupportedJsonCharVarcharMapKey(key, keyType))
} else {
None
}
}

// A VARCHAR(n) JSON object name is never trimmed: it must already be at most n characters.
private def checkVarcharJsonMapKey(key: UTF8String, keyType: VarcharType): Option[Throwable] = {
if (key.numChars() > keyType.length) {
Some(QueryExecutionErrors.unsupportedJsonCharVarcharMapKey(key, keyType))
} else {
None
}
}

/**
* Parse an object as a Array.
*/
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -2739,6 +2739,30 @@ private[sql] object QueryExecutionErrors extends QueryErrorsBase with ExecutionE
)
}

// SQLSTATE is 0A000 (feature-not-supported), not the 54006 of the value-side
// EXCEED_LIMIT_LENGTH: a too-long CHAR/VARCHAR value is truncatable overflow, but a JSON object
// name is a key that is never padded or trimmed, so an off-width name has no representation in
// a CHAR(n)/VARCHAR(n) map at all. It is a SparkRuntimeException (not
// SparkSQLFeatureNotSupportedException) because it is raised while parsing a row and must flow
// through Jackson's throw/catch parse-mode model (PERMISSIVE wraps it as a bad record, FAILFAST
// surfaces it).
def unsupportedJsonCharVarcharMapKey(
key: UTF8String, dataType: DataType): SparkRuntimeException = {
// The key is a raw JSON object name, which in map-shaped data is often user data (ids,
// emails) and can run up to Jackson's 50,000-character name limit, so cap what we echo.
val maxKeyChars = 128
val displayKey = if (key.numChars() > maxKeyChars) {
UTF8String.concat(key.substring(0, maxKeyChars), UTF8String.fromString("..."))
} else {
key
}
new SparkRuntimeException(
errorClass = "UNSUPPORTED_JSON_CHAR_VARCHAR_MAP_KEY",
messageParameters = Map(
"key" -> toSQLValue(displayKey, StringType),
"dataType" -> toSQLType(dataType)))
}

def timestampAddOverflowError(micros: Long, amount: Long, unit: String): ArithmeticException = {
new SparkArithmeticException(
errorClass = "DATETIME_OVERFLOW",
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -736,7 +736,7 @@ Project [to_date(26/October/2015, Some(dd/MMMMM/yyyy), Some(America/Los_Angeles)
-- !query
select from_json('{"d":"26/October/2015"}', 'd Date', map('dateFormat', 'dd/MMMMM/yyyy'))
-- !query analysis
Project [from_json(StructField(d,DateType,true), (dateFormat,dd/MMMMM/yyyy), {"d":"26/October/2015"}, Some(America/Los_Angeles), false) AS from_json({"d":"26/October/2015"})#x]
Project [from_json(StructField(d,DateType,true), (dateFormat,dd/MMMMM/yyyy), {"d":"26/October/2015"}, Some(America/Los_Angeles), false, false) AS from_json({"d":"26/October/2015"})#x]
+- OneRowRelation


Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -736,7 +736,7 @@ Project [to_date(26/October/2015, Some(dd/MMMMM/yyyy), Some(America/Los_Angeles)
-- !query
select from_json('{"d":"26/October/2015"}', 'd Date', map('dateFormat', 'dd/MMMMM/yyyy'))
-- !query analysis
Project [from_json(StructField(d,DateType,true), (dateFormat,dd/MMMMM/yyyy), {"d":"26/October/2015"}, Some(America/Los_Angeles), false) AS from_json({"d":"26/October/2015"})#x]
Project [from_json(StructField(d,DateType,true), (dateFormat,dd/MMMMM/yyyy), {"d":"26/October/2015"}, Some(America/Los_Angeles), false, false) AS from_json({"d":"26/October/2015"})#x]
+- OneRowRelation


Expand Down Expand Up @@ -1970,7 +1970,7 @@ Project [unix_timestamp(22 05 2020 Friday, dd MM yyyy EEEEE, Some(America/Los_An
-- !query
select from_json('{"t":"26/October/2015"}', 't Timestamp', map('timestampFormat', 'dd/MMMMM/yyyy'))
-- !query analysis
Project [from_json(StructField(t,TimestampType,true), (timestampFormat,dd/MMMMM/yyyy), {"t":"26/October/2015"}, Some(America/Los_Angeles), false) AS from_json({"t":"26/October/2015"})#x]
Project [from_json(StructField(t,TimestampType,true), (timestampFormat,dd/MMMMM/yyyy), {"t":"26/October/2015"}, Some(America/Los_Angeles), false, false) AS from_json({"t":"26/October/2015"})#x]
+- OneRowRelation


Expand Down
Loading