Skip to content

[SPARK-60146][PYTHON] Use len for MLlib vector dimensions - #59348

Draft
bearomorphism wants to merge 2 commits into
apache:masterfrom
bearomorphism:bearomorphism/fix-remove-type-ignore
Draft

bearomorphism wants to merge 2 commits into
apache:masterfrom
bearomorphism:bearomorphism/fix-remove-type-ignore

Conversation

@bearomorphism

@bearomorphism bearomorphism commented Oct 11, 2026 •

Copy link
Copy Markdown

What changes were proposed in this pull request?

Replace five .size accesses with len() in pyspark.mllib.classification, covering coefficient dimensions, multiclass prediction, and streaming initial weights. This removes four redundant type: ignore[attr-defined] comments.

The Vector base class and concrete implementations remain unchanged. Suppressions for undeclared methods such as dot() are outside this PR's scope.

Why are the changes needed?

Vector already declares __len__(), and existing dot-product and squared-distance dimension checks rely on it. Both built-in vector implementations return their logical dimension from len(), including trailing zeros in sparse vectors.

Using this existing interface avoids relying on an undeclared .size attribute or adding .size as a new requirement for all Vector subclasses.

Does this PR introduce any user-facing change?

No behavior change for the built-in DenseVector and SparseVector implementations. Custom subclasses used by these classification paths must implement the existing Vector.__len__() interface; providing only .size is no longer sufficient for these dimension accesses.

How was this patch tested?

Passed:

mypy --config-file python/mypy.ini --follow-imports=silent python/pyspark/mllib
ruff check python/pyspark/mllib/classification.py
ruff format --check python/pyspark/mllib/classification.py
git diff --check

Temporary Python checks passed 38 scenarios both before and after the revision, covering dense/sparse weights and features, multiclass prediction with and without bias, trailing zeros, and streaming initialization with empty, all-zero, and nonzero weights.

No permanent tests were added for this cleanup. No JVM-backed test suites were run.

Was this patch authored or co-authored using generative AI tooling?

Generated-by: Pi 1.1.0

@bearomorphism bearomorphism changed the title [MINOR][MLLIB] Declare size on PySpark Vector [MINOR][PYTHON] Declare size on PySpark Vector Oct 11, 2026
@bearomorphism bearomorphism changed the title [MINOR][PYTHON] Declare size on PySpark Vector [SPARK-60146][PYTHON] Declare size on PySpark Vector Oct 11, 2026
@bearomorphism
bearomorphism marked this pull request as ready for review October 11, 2026 05:17
@bearomorphism
bearomorphism marked this pull request as draft October 11, 2026 05:29
@bearomorphism bearomorphism changed the title [SPARK-60146][PYTHON] Declare size on PySpark Vector [SPARK-60146][PYTHON] Use len for MLlib vector dimensions Oct 11, 2026
@bearomorphism

Copy link
Copy Markdown
Author

After taking a second look I think we need a deeper discussion on the fix.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant