You are viewing a plain text version of this content. The canonical link for it is here.
Posted to reviews@spark.apache.org by GitBox <gi...@apache.org> on 2022/08/18 06:56:34 UTC
[GitHub] [spark] zhengruifeng opened a new pull request, #37564: [SPARK-40135][PS] Support ps.Index in DataFrame creation
zhengruifeng opened a new pull request, #37564:
URL: https://github.com/apache/spark/pull/37564
### What changes were proposed in this pull request?
Support `ps.Index` in DataFrame creation when `compute.ops_on_diff_frames` is `True`
### Why are the changes needed?
to support more options in DataFrame creation
### Does this PR introduce _any_ user-facing change?
Yes, `ps.Index` is supported
### How was this patch tested?
added UT
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r956530408
##########
python/pyspark/pandas/frame.py:
##########
@@ -375,6 +373,16 @@ class DataFrame(Frame, Generic[T]):
copy : boolean, default False
Copy data from inputs. Only affects DataFrame / 2d ndarray input
+ .. versionchanged:: 3.4.0
+ Since 3.4.0, it deals with `data` and `index` in this approach:
+ 1, when `data` is a distributed dataset (Internal DataFrame/Spark DataFrame/
+ pandas-on-Spark DataFrame/pandas-on-Spark Series), it will first parallize
+ the `index` if necessary, and then try to combine the `data` and `index`;
+ Note that in this case `compute.ops_on_diff_frames` should be turned on;
+ 2, when `data` is a local dataset (Pandas DataFrame/numpy ndarray/list/etc),
+ it will first collect the `index` to driver if necessary, and then apply
+ the `Pandas.DataFrame(...)` creation internally;
Review Comment:
done
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] itholic commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
itholic commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r955525708
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +420,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
Review Comment:
nit: Do we need to import `numpy` again and following examples ?? (and `pandas` as well ?)
##########
python/pyspark/pandas/tests/test_dataframe.py:
##########
@@ -108,13 +108,117 @@ def test_dataframe(self):
self.assert_eq(pdf, psdf)
def test_creation_index(self):
Review Comment:
Can we also have a test with various data types other than integer type ?? (string and timestamp ?)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support ps.Index in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r949701378
##########
python/pyspark/pandas/frame.py:
##########
@@ -455,10 +456,14 @@ def __init__( # type: ignore[no-untyped-def]
from pyspark.pandas.indexes.base import Index
if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
+ if get_option("compute.ops_on_diff_frames"):
+ index = index.to_pandas()
Review Comment:
I am still a lit confused:
do you mean:
1, parallelize data;
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] HyukjinKwon commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
HyukjinKwon commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r956853668
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +419,158 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=pd.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
a b c d e
- 0 3 1 4 9 8
- 1 4 8 4 8 4
- 2 7 6 5 6 7
- 3 8 7 9 1 0
- 4 2 5 4 3 9
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with pandas-on-Spark index:
Review Comment:
```suggestion
Constructing DataFrame from NumPy ndarray with pandas-on-Spark index:
```
##########
python/pyspark/pandas/frame.py:
##########
@@ -375,6 +373,16 @@ class DataFrame(Frame, Generic[T]):
copy : boolean, default False
Copy data from inputs. Only affects DataFrame / 2d ndarray input
+ .. versionchanged:: 3.4.0
+ Since 3.4.0, it deals with `data` and `index` in this approach:
+ 1, when `data` is a distributed dataset (Internal DataFrame/Spark DataFrame/
+ pandas-on-Spark DataFrame/pandas-on-Spark Series), it will first parallize
+ the `index` if necessary, and then try to combine the `data` and `index`;
+ Note that in this case `compute.ops_on_diff_frames` should be turned on;
+ 2, when `data` is a local dataset (Pandas DataFrame/numpy ndarray/list/etc),
+ it will first collect the `index` to driver if necessary, and then apply
+ the `Pandas.DataFrame(...)` creation internally;
Review Comment:
```suggestion
2. when `data` is a local dataset (pandas DataFrame, NumPy ndarray, list, etc.),
it will first collect the `index` to driver if necessary, and then apply
the `pandas.DataFrame(...)` creation internally;
```
##########
python/pyspark/pandas/frame.py:
##########
@@ -359,11 +359,9 @@ class DataFrame(Frame, Generic[T]):
Parameters
----------
- data : numpy ndarray (structured or homogeneous), dict, pandas DataFrame, Spark DataFrame \
- or pandas-on-Spark Series
+ data : numpy ndarray (structured or homogeneous), dict, pandas DataFrame,
Review Comment:
```suggestion
data : NumPy ndarray (structured or homogeneous), dict, pandas DataFrame,
```
##########
python/pyspark/pandas/frame.py:
##########
@@ -375,6 +373,16 @@ class DataFrame(Frame, Generic[T]):
copy : boolean, default False
Copy data from inputs. Only affects DataFrame / 2d ndarray input
+ .. versionchanged:: 3.4.0
+ Since 3.4.0, it deals with `data` and `index` in this approach:
+ 1, when `data` is a distributed dataset (Internal DataFrame/Spark DataFrame/
+ pandas-on-Spark DataFrame/pandas-on-Spark Series), it will first parallize
+ the `index` if necessary, and then try to combine the `data` and `index`;
+ Note that in this case `compute.ops_on_diff_frames` should be turned on;
Review Comment:
```suggestion
1. when `data` is a distributed dataset (internal DataFrame, PySpark DataFrame,
pandas-on-Spark DataFrame, and pandas-on-Spark Series), it will first parallelize
the `index` if necessary, and then try to combine the `data` and `index`;
Note that in this case `compute.ops_on_diff_frames` should be turned on;
```
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +419,158 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
Review Comment:
```suggestion
Constructing DataFrame from NumPy ndarray with pandas index:
```
##########
python/pyspark/pandas/frame.py:
##########
@@ -359,11 +359,9 @@ class DataFrame(Frame, Generic[T]):
Parameters
----------
- data : numpy ndarray (structured or homogeneous), dict, pandas DataFrame, Spark DataFrame \
- or pandas-on-Spark Series
+ data : numpy ndarray (structured or homogeneous), dict, pandas DataFrame,
+ Spark DataFrame, pandas-on-Spark DataFrame or pandas-on-Spark Series.
Review Comment:
```suggestion
PySpark DataFrame, pandas-on-Spark DataFrame or pandas-on-Spark Series.
```
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r956530421
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +420,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=pd.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
a b c d e
- 0 3 1 4 9 8
- 1 4 8 4 8 4
- 2 7 6 5 6 7
- 3 8 7 9 1 0
- 4 2 5 4 3 9
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=ps.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from Pandas DataFrame with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=pd.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Pandas DataFrame with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=ps.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Spark DataFrame with Pandas index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
+
+ Constructing DataFrame from Spark DataFrame with pandas-on-Spark index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
"""
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
+
+ pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
internal = InternalFrame.from_pandas(pdf)
+ index_assigned = True
+
+ if index is not None and not index_assigned:
+ data_df = ps.DataFrame(data=data, index=None, columns=columns, dtype=dtype, copy=copy)
+ index_ps = ps.Index(index)
+ index_df = index_ps.to_frame()
+
+ # drop un-matched rows in `data`
+ # note that `combine_frames` can not work with a MultiIndex for now
Review Comment:
done
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support ps.Index in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r949701378
##########
python/pyspark/pandas/frame.py:
##########
@@ -455,10 +456,14 @@ def __init__( # type: ignore[no-untyped-def]
from pyspark.pandas.indexes.base import Index
if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
+ if get_option("compute.ops_on_diff_frames"):
+ index = index.to_pandas()
Review Comment:
I am still a lit confused:
do you mean this?
1. parallelize data;
2. attach a `temp_index` with `AttachDistributedSequence` or `zipWithIndex` for both data and ps.Index;
3. inner join on the `temp_index`
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r955715479
##########
python/pyspark/pandas/tests/test_dataframe.py:
##########
@@ -108,13 +108,117 @@ def test_dataframe(self):
self.assert_eq(pdf, psdf)
def test_creation_index(self):
Review Comment:
Done
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng closed pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng closed pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
URL: https://github.com/apache/spark/pull/37564
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on pull request #37564: [SPARK-40135][PS][WIP] Support ps.Index in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on PR #37564:
URL: https://github.com/apache/spark/pull/37564#issuecomment-1221514409
one example:
```python
In [1]:
...: from pyspark.sql.types import *
...: from pyspark.sql import Column, DataFrame as SparkDataFrame, functions as F
...:
...: data = np.random.randn(5, 2)
...:
...: index = ps.Index([1, 3, 5, 7, 9])
...:
...: pdf = pd.DataFrame(data, columns=["A", "B"])
...:
...: ps.set_option("compute.ops_on_diff_frames", True)
...:
...: psdf = ps.DataFrame(data=pdf, index=index)
...:
In [2]: psdf
A B
NaN 1.283948 1.014323
1.0 -0.093965 -0.752795
NaN -1.254084 1.586583
3.0 1.133467 -0.773754
NaN -2.089362 2.568566
5.0 NaN NaN
7.0 NaN NaN
9.0 NaN NaN
```
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support ps.Index in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r949701378
##########
python/pyspark/pandas/frame.py:
##########
@@ -455,10 +456,14 @@ def __init__( # type: ignore[no-untyped-def]
from pyspark.pandas.indexes.base import Index
if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
+ if get_option("compute.ops_on_diff_frames"):
+ index = index.to_pandas()
Review Comment:
I am still a lit confused:
do you mean:
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r953622172
##########
python/pyspark/pandas/frame.py:
##########
@@ -425,42 +423,64 @@ class DataFrame(Frame, Generic[T]):
def __init__( # type: ignore[no-untyped-def]
Review Comment:
sure
##########
python/pyspark/pandas/frame.py:
##########
@@ -426,42 +426,82 @@ class DataFrame(Frame, Generic[T]):
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
- else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
+ if index is None:
+ data = data.to_frame()
+ internal = data._internal
+ elif isinstance(data, pd.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ internal = InternalFrame.from_pandas(pdf)
+ else:
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
+ if not isinstance(index, Index):
pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
- internal = InternalFrame.from_pandas(pdf)
+ internal = InternalFrame.from_pandas(pdf)
+ index_assigned = True
+
+ if index is not None and not index_assigned:
Review Comment:
cc @HyukjinKwon @ueshin PTAL, is this what you expected?
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +420,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=pd.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
a b c d e
- 0 3 1 4 9 8
- 1 4 8 4 8 4
- 2 7 6 5 6 7
- 3 8 7 9 1 0
- 4 2 5 4 3 9
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=ps.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from Pandas DataFrame with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=pd.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Pandas DataFrame with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=ps.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Spark DataFrame with Pandas index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
+
+ Constructing DataFrame from Spark DataFrame with pandas-on-Spark index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
"""
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
+
+ pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
internal = InternalFrame.from_pandas(pdf)
+ index_assigned = True
+
+ if index is not None and not index_assigned:
+ data_df = ps.DataFrame(data=data, index=None, columns=columns, dtype=dtype, copy=copy)
+ index_ps = ps.Index(index)
+ index_df = index_ps.to_frame()
+
+ # drop un-matched rows in `data`
+ # note that `combine_frames` can not work with a MultiIndex for now
Review Comment:
good idea
##########
python/pyspark/pandas/tests/test_dataframe.py:
##########
@@ -108,13 +108,117 @@ def test_dataframe(self):
self.assert_eq(pdf, psdf)
def test_creation_index(self):
Review Comment:
let me have a try
##########
python/pyspark/pandas/frame.py:
##########
@@ -375,6 +373,17 @@ class DataFrame(Frame, Generic[T]):
copy : boolean, default False
Copy data from inputs. Only affects DataFrame / 2d ndarray input
+ Notes
+ -----
+ Since 3.4.0, it deals with `index` in this way:
Review Comment:
good idea, will update!
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +420,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
Review Comment:
I think it's needed to be self-contained
##########
python/pyspark/pandas/frame.py:
##########
@@ -425,42 +423,64 @@ class DataFrame(Frame, Generic[T]):
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
+
+ pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
internal = InternalFrame.from_pandas(pdf)
+ index_assigned = True
+
+ if index is not None and not index_assigned:
+ data_df = ps.DataFrame(data=data, index=None, columns=columns, dtype=dtype, copy=copy)
+ index_ps = ps.Index(index)
+ index_df = index_ps.to_frame()
+
+ # drop un-matched rows in `data`
+ # note that `combine_frames` can not work with a MultiIndex for now
+ combined = combine_frames(data_df, index_df, how="right")
+ combined_labels = combined._internal.column_labels
+ index_labels = [label for label in combined_labels if label[0] == "that"]
+ combined = combined.set_index(index_labels)
+
+ combined._internal._column_labels = data_df._internal.column_labels
Review Comment:
I tried it, but it seem that we can not modify the level of labels in `copy`
while in this case, the label will be changed from `(this, A, )` to `(A,)`
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] HyukjinKwon commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
HyukjinKwon commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r953393376
##########
python/pyspark/pandas/frame.py:
##########
@@ -425,42 +423,64 @@ class DataFrame(Frame, Generic[T]):
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
+
+ pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
internal = InternalFrame.from_pandas(pdf)
+ index_assigned = True
+
+ if index is not None and not index_assigned:
+ data_df = ps.DataFrame(data=data, index=None, columns=columns, dtype=dtype, copy=copy)
+ index_ps = ps.Index(index)
+ index_df = index_ps.to_frame()
+
+ # drop un-matched rows in `data`
+ # note that `combine_frames` can not work with a MultiIndex for now
+ combined = combine_frames(data_df, index_df, how="right")
+ combined_labels = combined._internal.column_labels
+ index_labels = [label for label in combined_labels if label[0] == "that"]
+ combined = combined.set_index(index_labels)
+
+ combined._internal._column_labels = data_df._internal.column_labels
Review Comment:
Could we use `combined_internal.copy(...)` instead?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] itholic commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
itholic commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r955527398
##########
python/pyspark/pandas/tests/test_dataframe.py:
##########
@@ -108,13 +108,117 @@ def test_dataframe(self):
self.assert_eq(pdf, psdf)
def test_creation_index(self):
Review Comment:
Can we also have tests with various data types other than integer type ?? (string and timestamp ?)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r955714686
##########
python/pyspark/pandas/frame.py:
##########
@@ -425,42 +423,64 @@ class DataFrame(Frame, Generic[T]):
def __init__( # type: ignore[no-untyped-def]
Review Comment:
done
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support ps.Index in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r948975259
##########
python/pyspark/pandas/frame.py:
##########
@@ -455,10 +456,14 @@ def __init__( # type: ignore[no-untyped-def]
from pyspark.pandas.indexes.base import Index
if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
+ if get_option("compute.ops_on_diff_frames"):
+ index = index.to_pandas()
Review Comment:
oh, I misunderstood the task.
in this case, the data should be converted to `SparkDataFrame` or `InternalFrame ` and then joined with the ps.Index, right?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] ueshin commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
ueshin commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r956361716
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +420,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=pd.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
a b c d e
- 0 3 1 4 9 8
- 1 4 8 4 8 4
- 2 7 6 5 6 7
- 3 8 7 9 1 0
- 4 2 5 4 3 9
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=ps.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from Pandas DataFrame with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=pd.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Pandas DataFrame with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=ps.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Spark DataFrame with Pandas index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
+
+ Constructing DataFrame from Spark DataFrame with pandas-on-Spark index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
"""
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
+
+ pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
internal = InternalFrame.from_pandas(pdf)
+ index_assigned = True
+
+ if index is not None and not index_assigned:
+ data_df = ps.DataFrame(data=data, index=None, columns=columns, dtype=dtype, copy=copy)
+ index_ps = ps.Index(index)
+ index_df = index_ps.to_frame()
+
+ # drop un-matched rows in `data`
+ # note that `combine_frames` can not work with a MultiIndex for now
Review Comment:
Could you also update the comment here as @itholic mentions?
Also could you raise an error with an appropriate error message for the case?
##########
python/pyspark/pandas/frame.py:
##########
@@ -375,6 +373,16 @@ class DataFrame(Frame, Generic[T]):
copy : boolean, default False
Copy data from inputs. Only affects DataFrame / 2d ndarray input
+ .. versionchanged:: 3.4.0
+ Since 3.4.0, it deals with `data` and `index` in this approach:
+ 1, when `data` is a distributed dataset (Internal DataFrame/Spark DataFrame/
+ pandas-on-Spark DataFrame/pandas-on-Spark Series), it will first parallize
+ the `index` if necessary, and then try to combine the `data` and `index`;
+ Note that in this case `compute.ops_on_diff_frames` should be turned on;
+ 2, when `data` is a local dataset (Pandas DataFrame/numpy ndarray/list/etc),
+ it will first collect the `index` to driver if necessary, and then apply
+ the `Pandas.DataFrame(...)` creation internally;
Review Comment:
I guess we need indent for the comment of `versionchanged`?
```py
.. versionchanged:: 3.4.0
Since 3.4.0, ...
```
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +419,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=pd.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
a b c d e
- 0 3 1 4 9 8
- 1 4 8 4 8 4
- 2 7 6 5 6 7
- 3 8 7 9 1 0
- 4 2 5 4 3 9
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=ps.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from Pandas DataFrame with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=pd.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Pandas DataFrame with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=ps.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Spark DataFrame with Pandas index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
+
+ Constructing DataFrame from Spark DataFrame with pandas-on-Spark index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
"""
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
Review Comment:
This should be `_to_pandas()` to avoid warnings?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] itholic commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
itholic commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r955525103
##########
python/pyspark/pandas/frame.py:
##########
@@ -375,6 +373,17 @@ class DataFrame(Frame, Generic[T]):
copy : boolean, default False
Copy data from inputs. Only affects DataFrame / 2d ndarray input
+ Notes
+ -----
+ Since 3.4.0, it deals with `index` in this way:
Review Comment:
Can we use `.. versionchanged:: 3.4.0` instead if it describes about the different behavior with previous version ??
##########
python/pyspark/pandas/tests/test_dataframe.py:
##########
@@ -108,13 +108,117 @@ def test_dataframe(self):
self.assert_eq(pdf, psdf)
def test_creation_index(self):
Review Comment:
Can we also test with various data types other than integer type ?? (string and timestamp ?)
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +420,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=pd.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
a b c d e
- 0 3 1 4 9 8
- 1 4 8 4 8 4
- 2 7 6 5 6 7
- 3 8 7 9 1 0
- 4 2 5 4 3 9
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=ps.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from Pandas DataFrame with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=pd.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Pandas DataFrame with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=ps.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Spark DataFrame with Pandas index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
+
+ Constructing DataFrame from Spark DataFrame with pandas-on-Spark index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
"""
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
+
+ pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
internal = InternalFrame.from_pandas(pdf)
+ index_assigned = True
+
+ if index is not None and not index_assigned:
+ data_df = ps.DataFrame(data=data, index=None, columns=columns, dtype=dtype, copy=copy)
+ index_ps = ps.Index(index)
+ index_df = index_ps.to_frame()
+
+ # drop un-matched rows in `data`
+ # note that `combine_frames` can not work with a MultiIndex for now
Review Comment:
qq: Maybe should we create a JIRA and make it TODO such as `TODO(SPARK-XXXXX): Support MultiIndex` ??
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +420,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
Review Comment:
nit: Do we need to import `numpy` here and following examples ?? (and `pandas` as well ?)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r955715324
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +420,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=pd.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
a b c d e
- 0 3 1 4 9 8
- 1 4 8 4 8 4
- 2 7 6 5 6 7
- 3 8 7 9 1 0
- 4 2 5 4 3 9
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=ps.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from Pandas DataFrame with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=pd.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Pandas DataFrame with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=ps.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Spark DataFrame with Pandas index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
+
+ Constructing DataFrame from Spark DataFrame with pandas-on-Spark index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
"""
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
+
+ pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
internal = InternalFrame.from_pandas(pdf)
+ index_assigned = True
+
+ if index is not None and not index_assigned:
+ data_df = ps.DataFrame(data=data, index=None, columns=columns, dtype=dtype, copy=copy)
+ index_ps = ps.Index(index)
+ index_df = index_ps.to_frame()
+
+ # drop un-matched rows in `data`
+ # note that `combine_frames` can not work with a MultiIndex for now
Review Comment:
I file a ticke https://issues.apache.org/jira/browse/SPARK-40226 to track this.
thanks
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] itholic commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
itholic commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r956855212
##########
python/pyspark/pandas/tests/test_dataframe.py:
##########
@@ -108,13 +108,117 @@ def test_dataframe(self):
self.assert_eq(pdf, psdf)
def test_creation_index(self):
Review Comment:
Nice 👍
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on PR #37564:
URL: https://github.com/apache/spark/pull/37564#issuecomment-1231184909
Merged into master, thank you @itholic @HyukjinKwon @ueshin so much!
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on PR #37564:
URL: https://github.com/apache/spark/pull/37564#issuecomment-1222125737
cc @ueshin @HyukjinKwon @xinrong-meng This PR is ready to review. Thanks!
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] HyukjinKwon commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
HyukjinKwon commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r953393776
##########
python/pyspark/pandas/frame.py:
##########
@@ -425,42 +423,64 @@ class DataFrame(Frame, Generic[T]):
def __init__( # type: ignore[no-untyped-def]
Review Comment:
BTW, should we maybe add some examples above in the docstring?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r956530355
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +419,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=pd.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
a b c d e
- 0 3 1 4 9 8
- 1 4 8 4 8 4
- 2 7 6 5 6 7
- 3 8 7 9 1 0
- 4 2 5 4 3 9
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=ps.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from Pandas DataFrame with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=pd.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Pandas DataFrame with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=ps.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Spark DataFrame with Pandas index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
+
+ Constructing DataFrame from Spark DataFrame with pandas-on-Spark index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
"""
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
Review Comment:
done
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] itholic commented on a diff in pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
itholic commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r956854955
##########
python/pyspark/pandas/frame.py:
##########
@@ -411,56 +420,154 @@ class DataFrame(Frame, Generic[T]):
Constructing DataFrame from numpy ndarray:
- >>> df2 = ps.DataFrame(np.random.randint(low=0, high=10, size=(5, 5)),
- ... columns=['a', 'b', 'c', 'd', 'e'])
- >>> df2 # doctest: +SKIP
+ >>> import numpy as np
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 0 1 2 3 4 5
+ 1 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=pd.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
a b c d e
- 0 3 1 4 9 8
- 1 4 8 4 8 4
- 2 7 6 5 6 7
- 3 8 7 9 1 0
- 4 2 5 4 3 9
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from numpy ndarray with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> ps.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... index=ps.Index([1, 4]), columns=['a', 'b', 'c', 'd', 'e'])
+ a b c d e
+ 1 1 2 3 4 5
+ 4 6 7 8 9 0
+
+ Constructing DataFrame from Pandas DataFrame with Pandas index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=pd.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Pandas DataFrame with pandas-on-Spark index:
+
+ >>> import numpy as np
+ >>> import pandas as pd
+ >>> pdf = pd.DataFrame(data=np.array([[1, 2, 3, 4, 5], [6, 7, 8, 9, 0]]),
+ ... columns=['a', 'b', 'c', 'd', 'e'])
+ >>> ps.DataFrame(data=pdf, index=ps.Index([1, 4]))
+ a b c d e
+ 1 6.0 7.0 8.0 9.0 0.0
+ 4 NaN NaN NaN NaN NaN
+
+ Constructing DataFrame from Spark DataFrame with Pandas index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=pd.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
+
+ Constructing DataFrame from Spark DataFrame with pandas-on-Spark index:
+
+ >>> import pandas as pd
+ >>> sdf = spark.createDataFrame([("Data", 1), ("Bricks", 2)], ["x", "y"])
+ >>> ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ Traceback (most recent call last):
+ ...
+ ValueError: Cannot combine the series or dataframe...'compute.ops_on_diff_frames' option.
+
+ Need to enable 'compute.ops_on_diff_frames' to combine SparkDataFrame and Pandas index
+
+ >>> with ps.option_context("compute.ops_on_diff_frames", True):
+ ... ps.DataFrame(data=sdf, index=ps.Index([0, 1, 2]))
+ x y
+ 0 Data 1.0
+ 1 Bricks 2.0
+ 2 None NaN
"""
def __init__( # type: ignore[no-untyped-def]
self, data=None, index=None, columns=None, dtype=None, copy=False
):
+ index_assigned = False
if isinstance(data, InternalFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = data
+ if index is None:
+ internal = data
elif isinstance(data, SparkDataFrame):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ if index is None:
+ internal = InternalFrame(spark_frame=data, index_spark_columns=None)
+ elif isinstance(data, ps.DataFrame):
+ assert columns is None
+ assert dtype is None
+ assert not copy
+ if index is None:
+ internal = data._internal.resolved_copy
elif isinstance(data, ps.Series):
- assert index is None
assert columns is None
assert dtype is None
assert not copy
- data = data.to_frame()
- internal = data._internal
+ if index is None:
+ internal = data.to_frame()._internal.resolved_copy
else:
- if isinstance(data, pd.DataFrame):
- assert index is None
- assert columns is None
- assert dtype is None
- assert not copy
- pdf = data
- else:
- from pyspark.pandas.indexes.base import Index
+ from pyspark.pandas.indexes.base import Index
- if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
- pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
+ if index is not None and isinstance(index, Index):
+ # with local data, collect ps.Index to driver
+ # to avoid mismatched results between
+ # ps.DataFrame([1, 2], index=ps.Index([1, 2]))
+ # and
+ # pd.DataFrame([1, 2], index=pd.Index([1, 2]))
+ index = index.to_pandas()
+
+ pdf = pd.DataFrame(data=data, index=index, columns=columns, dtype=dtype, copy=copy)
internal = InternalFrame.from_pandas(pdf)
+ index_assigned = True
+
+ if index is not None and not index_assigned:
+ data_df = ps.DataFrame(data=data, index=None, columns=columns, dtype=dtype, copy=copy)
+ index_ps = ps.Index(index)
+ index_df = index_ps.to_frame()
+
+ # drop un-matched rows in `data`
+ # note that `combine_frames` can not work with a MultiIndex for now
Review Comment:
Nice!
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] itholic commented on pull request #37564: [SPARK-40135][PS] Support `data` mixed with `index` in DataFrame creation
Posted by GitBox <gi...@apache.org>.
itholic commented on PR #37564:
URL: https://github.com/apache/spark/pull/37564#issuecomment-1229693201
Looks pretty good to me once https://github.com/apache/spark/pull/37564#pullrequestreview-1088007634 is resolved.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] zhengruifeng commented on pull request #37564: [SPARK-40135][PS] Support ps.Index in DataFrame creation
Posted by GitBox <gi...@apache.org>.
zhengruifeng commented on PR #37564:
URL: https://github.com/apache/spark/pull/37564#issuecomment-1219191446
cc @HyukjinKwon @ueshin @xinrong-meng
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org
[GitHub] [spark] HyukjinKwon commented on a diff in pull request #37564: [SPARK-40135][PS] Support ps.Index in DataFrame creation
Posted by GitBox <gi...@apache.org>.
HyukjinKwon commented on code in PR #37564:
URL: https://github.com/apache/spark/pull/37564#discussion_r948918084
##########
python/pyspark/pandas/frame.py:
##########
@@ -455,10 +456,14 @@ def __init__( # type: ignore[no-untyped-def]
from pyspark.pandas.indexes.base import Index
if isinstance(index, Index):
- raise TypeError(
- "The given index cannot be a pandas-on-Spark index. "
- "Try pandas index or array-like."
- )
+ if get_option("compute.ops_on_diff_frames"):
+ index = index.to_pandas()
Review Comment:
I think we should probably create pandas-on-Spark DataFrame with joining (e.g., df1 + df2 when `compute.ops_on_diff_frames` is on) instead of collecting it to the driver side.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org