SeriesGroupBy.nunique raises IndexError for empty Series with categorical variables #21334

parasurama · 2018-06-05T23:52:46Z

Code Sample

df = pd.DataFrame({'id': ['1001', '1001', '1002', '1002', '1003', '1004', '1005', '1006'],
                   'cat': ['A', 'B', 'C', 'C', 'C', 'B', 'B', 'D'],
                   'cat2': ['a', 'a', 'b', 'b', 'c', 'd', 'b', 'd']})
df['cat'] = df['cat'].astype('category')
df['cat2'] = df['cat2'].astype('category')

Problem description

If I subset this dataframe so groupby will return some empty Series and apply the following lambda function, it throws an index error IndexError: index 0 is out of bounds for axis 1 with size 0

df2 = df[df['cat2'].isin(['a', 'b'])]
df2.groupby('cat').apply(lambda x: x.groupby('cat2')['id'].agg('nunique'))  # throws IndexError

Note that df2.groupby(['cat', 'cat2'])['id'].agg('nunique') doesn't throw an error.
However, I have .groupby('cat2')['id'].agg('nunique') as part of a separate function. Hence using lambda for illustrative purposes.
Also note that the count aggregator doesn't throw an error:
df2.groupby('cat').apply(lambda x: x.groupby('cat2')['id'].agg('count'))
If i don't categorize the variables, it doesn't throw an error

Expected Output

cat cat2
A a 1
B a 1
B b 1
C b 1

Output of `pd.show_versions()`

[paste the output of `pd.show_versions()` here below this line]
INSTALLED VERSIONS

commit: None
python: 3.6.1.final.0
python-bits: 64
OS: Darwin
OS-release: 17.2.0
machine: x86_64
processor: i386
byteorder: little
LC_ALL: None
LANG: None
LOCALE: en_US.UTF-8
pandas: 0.22.0
pytest: 3.2.3
pip: 10.0.1
setuptools: 39.1.0
Cython: 0.25.2
numpy: 1.14.2
scipy: 1.0.1
pyarrow: None
xarray: None
IPython: 5.3.0
sphinx: 1.5.6
patsy: 0.4.1
dateutil: 2.7.2
pytz: 2018.3
blosc: None
bottleneck: 1.2.1
tables: 3.3.0
numexpr: 2.6.2
feather: None
matplotlib: 2.0.2
openpyxl: 2.4.7
xlrd: 1.0.0
xlwt: 1.2.0
xlsxwriter: 0.9.6
lxml: 3.7.3
bs4: 4.6.0
html5lib: 1.0b8
sqlalchemy: 1.1.9
pymysql: None
psycopg2: 2.7.1 (dt dec pq3 ext lo64)
jinja2: 2.9.6
s3fs: None
fastparquet: None
pandas_gbq: None
pandas_datareader: None

The text was updated successfully, but these errors were encountered:

gfyoung · 2018-06-06T07:14:08Z

Admittedly, the code sample is a little unusual, but nevertheless, IMO it should work whether the variables are categorical or not.

@jreback : thoughts?

rhshadrach · 2022-12-06T02:02:24Z

Here is a simplified example; it happens on any empty categoricals.

ser = pd.Series([1]).astype('category')[:0]
gb = ser.groupby(ser)
result = gb.nunique()
print(result)

gfyoung added Groupby Categorical Categorical Data Type labels Jun 6, 2018

gfyoung mentioned this issue Jun 8, 2018

groupby on 2 categorical columns, when one categorical is based on datetimes, incorrectly returns all NaN dataframe #21390

Closed

mroeschke added the Bug label Jun 28, 2020

rhshadrach self-assigned this Dec 6, 2022

rhshadrach mentioned this issue Dec 6, 2022

BUG: SeriesGroupBy.nunique raises when grouping with an empty categrorical #50081

Merged

5 tasks

mroeschke closed this as completed in #50081 Dec 8, 2022

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

SeriesGroupBy.nunique raises IndexError for empty Series with categorical variables #21334

SeriesGroupBy.nunique raises IndexError for empty Series with categorical variables #21334

parasurama commented Jun 5, 2018 •

edited

Loading

[paste the output of `pd.show_versions()` here below this line]
INSTALLED VERSIONS

gfyoung commented Jun 6, 2018 •

edited

Loading

rhshadrach commented Dec 6, 2022

SeriesGroupBy.nunique raises IndexError for empty Series with categorical variables #21334

SeriesGroupBy.nunique raises IndexError for empty Series with categorical variables #21334

Comments

parasurama commented Jun 5, 2018 • edited Loading

Code Sample

Problem description

Expected Output

Output of pd.show_versions()

[paste the output of pd.show_versions() here below this line] INSTALLED VERSIONS

gfyoung commented Jun 6, 2018 • edited Loading

rhshadrach commented Dec 6, 2022

parasurama commented Jun 5, 2018 •

edited

Loading

Output of `pd.show_versions()`

[paste the output of `pd.show_versions()` here below this line]
INSTALLED VERSIONS

gfyoung commented Jun 6, 2018 •

edited

Loading