Repository navigation
Commit be37962
[SPARK-59249][SQL][FOLLOWUP] Include SQL type in Python UDT identity
### What changes were proposed in this pull request?
This is a follow-up to [SPARK-59249](https://issues.apache.org/jira/browse/SPARK-59249) and [`this commit`](ebb0bd2), which routed grouped-key ordering through `InternalRowComparableWrapper`'s shared ordering cache.
This PR includes both `pyUDT` and `sqlType` in `PythonUserDefinedType.equals` and `hashCode`.
`serializedPyClass` remains excluded from identity, and `acceptsType` retains its existing compatibility semantics.
The PR also adds regression coverage for the shared `InternalRowComparableWrapper` ordering cache in both lookup orders, plus focused equality and hash-code tests.
### Why are the changes needed?
`InternalRowComparableWrapper` caches generated orderings by data type. Python UDT equality previously considered only the Python class name, so UDTs with the same class name but different underlying SQL types could share a cache entry.
For example, binary and `UTF8_LCASE` string-backed UDTs could reuse the first generated comparator, producing lookup-order-dependent collation semantics.
### How was this patch tested?
Added regression tests to `InternalRowComparableWrapperSuite`, covering binary-first and UTF8_LCASE-first cache lookup orders, and to `DataTypeSuite`, covering equality, hash codes, differing SQL types, and exclusion of serialized Python class data.
Ran:
```
build/sbt \
'catalyst/testOnly org.apache.spark.sql.catalyst.util.InternalRowComparableWrapperSuite' \
'catalyst/testOnly org.apache.spark.sql.types.DataTypeSuite'
```
`InternalRowComparableWrapperSuite`: 7 passed; `DataTypeSuite`: 352 passed.
### Was this patch authored or co-authored using generative AI tooling?
Yes co-authored-by: OpenAI Codex GPT-5.6 sol.
Closes #58799 from lavanv11/sc-244639-python-udt-ordering.
Authored-by: Lavan Vivekanandasarma <lavan.vivek@databricks.com>
Signed-off-by: Yicong-Huang <17627829+Yicong-Huang@users.noreply.github.com>1 parent 169da93 commit be37962
3 files changed
Lines changed: 46 additions & 2 deletions
File tree
- sql
- api/src/main/scala/org/apache/spark/sql/types
- catalyst/src/test/scala/org/apache/spark/sql
- catalyst/util
- types
Lines changed: 2 additions & 2 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
154 | 154 | | |
155 | 155 | | |
156 | 156 | | |
157 | | - | |
| 157 | + | |
158 | 158 | | |
159 | 159 | | |
160 | 160 | | |
161 | | - | |
| 161 | + | |
162 | 162 | | |
Lines changed: 32 additions & 0 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
17 | 17 | | |
18 | 18 | | |
19 | 19 | | |
| 20 | + | |
| 21 | + | |
20 | 22 | | |
21 | 23 | | |
22 | 24 | | |
23 | 25 | | |
| 26 | + | |
24 | 27 | | |
25 | 28 | | |
26 | 29 | | |
27 | 30 | | |
28 | 31 | | |
29 | 32 | | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
30 | 54 | | |
31 | 55 | | |
32 | 56 | | |
| |||
95 | 119 | | |
96 | 120 | | |
97 | 121 | | |
| 122 | + | |
| 123 | + | |
| 124 | + | |
| 125 | + | |
| 126 | + | |
| 127 | + | |
| 128 | + | |
| 129 | + | |
98 | 130 | | |
99 | 131 | | |
100 | 132 | | |
| |||
Lines changed: 12 additions & 0 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
38 | 38 | | |
39 | 39 | | |
40 | 40 | | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
41 | 53 | | |
42 | 54 | | |
43 | 55 | | |
| |||
0 commit comments