-
Notifications
You must be signed in to change notification settings - Fork 29.3k
Pull requests: apache/spark
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[SPARK-58518][SQL] Do not duplicate input paths when globbing is disabled
#57769
opened Aug 4, 2026 by
Joorgem
Loading…
[SPARK-57370][SQL][FOLLOWUP] Route codegen referencing an unnameable class to Janino when narrowing is unsound
#57768
opened Aug 4, 2026 by
LuciferYang
Contributor
•
Draft
[SPARK-58567][SQL] Use nonEmpty/isEmpty instead of length comparisons in StateDataSource
#57767
opened Aug 4, 2026 by
uros-b
Member
Loading…
[SPARK-58566][SQL] Add tests for QuotingUtils.quoteIdentifier and escapeSingleQuotedString
#57766
opened Aug 4, 2026 by
uros-b
Member
Loading…
[SPARK-58562][SQL][DOCS] Document subplan merging in the SQL performance tuning guide
#57765
opened Aug 4, 2026 by
peter-toth
Contributor
Loading…
[SPARK-58561][INFRA] Support backporting a PR merged into a non-default branch
#57764
opened Aug 4, 2026 by
uros-b
Member
Loading…
[SPARK-58559][PYTHON] Package all of sbin in PySpark classic distribution
#57763
opened Aug 4, 2026 by
nchammas
Contributor
Loading…
[SPARK-58558][SQL] Make requireAllClusterKeysForCoPartition check key coverage instead of exact match for SPJ
#57762
opened Aug 4, 2026 by
pan3793
Member
Loading…
[WIP][ML] Avoid duplicate StringIndexer skip lookups
#57761
opened Aug 4, 2026 by
zhengruifeng
Contributor
•
Draft
[SPARK-58207][SQL][FOLLOWUP] Skip runtime filter pushdown for nondeterministic filters
#57760
opened Aug 4, 2026 by
peter-toth
Contributor
Loading…
[SPARK-49828][SQL] Make Column(expression) usable outside of the org.apache.spark package
#57759
opened Aug 4, 2026 by
peter-toth
Contributor
Loading…
[SPARK-58553][PS] Use native Spark functions for NumPy fmax and fmin
#57758
opened Aug 4, 2026 by
zhengruifeng
Contributor
Loading…
[SPARK-58557][ML] Reuse ML vector data type singleton
#57757
opened Aug 4, 2026 by
zhengruifeng
Contributor
Loading…
[WIP][MLLIB] Rewrite CountVectorizer fitting with DataFrame APIs
#57755
opened Aug 4, 2026 by
zhengruifeng
Contributor
•
Draft
[SPARK-58556][CONNECT][UI] Show ML cache status in Spark Connect UI
#57754
opened Aug 4, 2026 by
zhengruifeng
Contributor
•
Draft
[SPARK-58549][SQL] Preserve key-grouped partitioning and ordering across a DSv2 scan merge
#57753
opened Aug 4, 2026 by
peter-toth
Contributor
Loading…
[SPARK-58551][PYTHON] Python Data Sources Limit Pushdown API
#57752
opened Aug 4, 2026 by
ganeshashree
Contributor
Loading…
[SPARK-58552][SQL][UI] Add total task time column to the SQL / DataFrame tab
#57751
opened Aug 4, 2026 by
ulysses-you
Contributor
Loading…
[SPARK-58550][ML] Delay GaussianMixture aggregation allocations
#57750
opened Aug 4, 2026 by
zhengruifeng
Contributor
•
Draft
[SPARK-58548][PS] Use native Spark function for NumPy heaviside
#57749
opened Aug 4, 2026 by
zhengruifeng
Contributor
Loading…
[SPARK-36284][CORE][SHUFFLE] Add shuffle checksum support for push-based shuffle
#57748
opened Aug 4, 2026 by
Dreamstick9
Loading…
[SPARK-58547][CONNECT] Expose operation IDs for end-to-end request attribution
#57747
opened Aug 4, 2026 by
cloud-fan
Contributor
Loading…
[SPARK-58544][SQL] Fix vector distance and norm functions returning wrong results from intermediate float overflow
#57746
opened Aug 4, 2026 by
SEPURI-SAI-KRISHNA
Loading…
[SPARK-58511][SQL] Bypass ineffective pre-shuffle partial aggregation at runtime
#57742
opened Aug 4, 2026 by
ulysses-you
Contributor
Loading…
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-07-04.