Skip to content

Commit 71f1edc

Browse files
ueshinHyukjinKwon
authored andcommitted
[SPARK-56016][PS] Preserve named Series columns in concat with ignore_index on pandas 3
### What changes were proposed in this pull request? This PR updates pandas-on-Spark `concat` for row-wise concatenation of mixed `DataFrame` and `Series` inputs when `ignore_index=True`. In pandas 2.x, pandas treats the `Series` input as an anonymous column named `0` in this case. In pandas 3.0+, pandas preserves the `Series` name and aligns it with matching `DataFrame` columns instead. pandas-on-Spark was still always converting mixed `Series` inputs to column `0`, which caused `pyspark.pandas.tests.test_namespace.NamespaceTests.test_concat_index_axis` to fail under pandas 3. To match pandas behavior, this patch keeps the existing behavior for pandas 2.x, but preserves the `Series` name for mixed `DataFrame`/`Series` concatenation on pandas 3.0+. This PR also updates the related namespace test to use `subTest(...)` for each concat case so the failing combination is easier to identify. ### Why are the changes needed? Without this change, pandas-on-Spark produces incorrect columns for mixed `DataFrame`/`Series` concatenation with `ignore_index=True` on pandas 3.0+. For example, concatenating a `DataFrame` with a named `Series` should reuse the `Series` name as the output column on pandas 3, but pandas-on-Spark instead creates a separate column `0`. That causes behavior differences from pandas and breaks the existing namespace concat test in the pandas 3 environment. ### Does this PR introduce _any_ user-facing change? Yes, it will behave more like pandas 3. ### How was this patch tested? The existing tests should pass. ### Was this patch authored or co-authored using generative AI tooling? Generated-by: OpenAI Codex (GPT-5) Closes #54837 from ueshin/issues/SPARK-56016/concat. Authored-by: Takuya Ueshin <ueshin@databricks.com> Signed-off-by: Hyukjin Kwon <gurwls223@apache.org>
1 parent 6f3ece6 commit 71f1edc

2 files changed

Lines changed: 17 additions & 9 deletions

File tree

‎python/pyspark/pandas/namespace.py‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2655,6 +2655,10 @@ def resolve_func(psdf, this_column_labels, that_column_labels):
26552655
series_names.add(obj.name)
26562656
if not ignore_index and not should_return_series:
26572657
new_objs.append(obj.to_frame())
2658+
elif LooseVersion(pd.__version__) >= "3.0.0" and not should_return_series:
2659+
# pandas 3 preserves a named Series as its own column during
2660+
# row-wise concat with ignore_index=True instead of renaming it to 0.
2661+
new_objs.append(obj.to_frame())
26582662
else:
26592663
new_objs.append(obj.to_frame(DEFAULT_SERIES_NAME))
26602664
else:

‎python/pyspark/pandas/tests/test_namespace.py‎

Lines changed: 13 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -361,17 +361,21 @@ def test_concat_index_axis(self):
361361
([psdf, psdf["C"]], [pdf, pdf["C"]]),
362362
([psdf["C"], psdf], [pdf["C"], pdf]),
363363
]
364-
for psdfs, pdfs in series_objs:
365-
for ignore_index, join, sort in itertools.product(ignore_indexes, joins, sorts):
366-
self.assert_eq(
367-
ps.concat(psdfs, ignore_index=ignore_index, join=join, sort=sort),
368-
pd.concat(pdfs, ignore_index=ignore_index, join=join, sort=sort),
369-
)
364+
365+
for ignore_index, join, sort in itertools.product(ignore_indexes, joins, sorts):
366+
for i, (psdfs, pdfs) in enumerate(series_objs):
367+
with self.subTest(
368+
ignore_index=ignore_index, join=join, sort=sort, pair=i, index="single"
369+
):
370+
self.assert_eq(
371+
ps.concat(psdfs, ignore_index=ignore_index, join=join, sort=sort),
372+
pd.concat(pdfs, ignore_index=ignore_index, join=join, sort=sort),
373+
)
370374

371375
for ignore_index, join, sort in itertools.product(ignore_indexes, joins, sorts):
372376
for i, (psdfs, pdfs) in enumerate(objs):
373377
with self.subTest(
374-
ignore_index=ignore_index, join=join, sort=sort, pdfs=pdfs, pair=i
378+
ignore_index=ignore_index, join=join, sort=sort, pair=i, index="single"
375379
):
376380
self.assert_eq(
377381
ps.concat(psdfs, ignore_index=ignore_index, join=join, sort=sort),
@@ -413,7 +417,7 @@ def test_concat_index_axis(self):
413417
for ignore_index, sort in itertools.product(ignore_indexes, sorts):
414418
for i, (psdfs, pdfs) in enumerate(objs):
415419
with self.subTest(
416-
ignore_index=ignore_index, join="outer", sort=sort, pdfs=pdfs, pair=i
420+
ignore_index=ignore_index, join="outer", sort=sort, pair=i, index="multi"
417421
):
418422
self.assert_eq(
419423
ps.concat(psdfs, ignore_index=ignore_index, join="outer", sort=sort),
@@ -424,7 +428,7 @@ def test_concat_index_axis(self):
424428
for ignore_index in ignore_indexes:
425429
for i, (psdfs, pdfs) in enumerate(objs):
426430
with self.subTest(
427-
ignore_index=ignore_index, join="inner", sort=True, pdfs=pdfs, pair=i
431+
ignore_index=ignore_index, join="inner", sort=True, pair=i, index="multi"
428432
):
429433
self.assert_eq(
430434
ps.concat(psdfs, ignore_index=ignore_index, join="inner", sort=True),

0 commit comments

Comments
 (0)