drop_duplicates slow for data frame of boolean columns #12963

tdhopper · 2016-04-22T21:05:36Z

In trying to find out the fast way to count unique rows in a data frame, I stumbled upon an issue where dropping duplicates from a data frame of booleans is quite slow. @jreback asked me to open a ticket.

In [1]: import pandas as pd

In [2]: import numpy as np

In [3]: N = 100000

In [4]: np.random.seed(1234)

In [5]: df = pd.DataFrame({str(i) : [bool(x) for x in np.random.randint(0,2,size=N)] for i in range(500)})

In [6]: %timeit df.drop_duplicates()
1 loop, best of 3: 3.01 s per loop

In [7]: %timeit df.astype(int).drop_duplicates()
1 loop, best of 3: 1.4 s per loop

INSTALLED VERSIONS
------------------
commit: None
python: 3.5.1.final.0
python-bits: 64
OS: Darwin
OS-release: 15.4.0
machine: x86_64
processor: i386
byteorder: little
LC_ALL: None
LANG: en_US.UTF-8

pandas: 0.18.0
nose: 1.3.7
pip: 8.1.1
setuptools: 20.3
Cython: 0.23.4
numpy: 1.10.4
scipy: 0.17.0
statsmodels: 0.6.1
xarray: None
IPython: 4.1.2
sphinx: 1.3.5
patsy: 0.4.0
dateutil: 2.5.1
pytz: 2016.2
blosc: None
bottleneck: 1.0.0
tables: 3.2.2
numexpr: 2.5
matplotlib: 1.5.1
openpyxl: 2.3.2
xlrd: 0.9.4
xlwt: 1.0.0
xlsxwriter: 0.8.4
lxml: 3.6.0
bs4: 4.4.1
html5lib: None
httplib2: None
apiclient: None
sqlalchemy: 1.0.12
pymysql: None
psycopg2: None
jinja2: 2.8
boto: 2.39.0

The text was updated successfully, but these errors were encountered:

sinhrks · 2016-04-25T00:18:07Z

xref #10235

Add whatsnew

Add whatsnew Add dtype label and reorg logic

closes pandas-dev#12963 Author: Matt Roeschke <[email protected]> Closes pandas-dev#15738 from mroeschke/fix_12963 and squashes the following commits: a020c10 [Matt Roeschke] PERF: Improve drop_duplicates for bool columns (pandas-dev#12963)

jreback added Performance Memory or execution speed performance Dtype Conversions Unexpected or buggy dtype conversions labels Apr 22, 2016

jreback added this to the 0.18.2 milestone Apr 22, 2016

jorisvandenbossche modified the milestones: Next Major Release, 0.19.0 Aug 13, 2016

mroeschke added a commit to mroeschke/pandas that referenced this issue Mar 20, 2017

PERF: Improve drop_duplicates for bool columns (pandas-dev#12963)

9334282

Add whatsnew

mroeschke mentioned this issue Mar 20, 2017

PERF: Improve drop_duplicates for bool columns (#12963) #15738

Closed

4 tasks

jreback modified the milestones: 0.20.0, Next Major Release Mar 20, 2017

mroeschke added a commit to mroeschke/pandas that referenced this issue Mar 20, 2017

PERF: Improve drop_duplicates for bool columns (pandas-dev#12963)

a020c10

Add whatsnew Add dtype label and reorg logic

jreback closed this as completed in f2e942e Mar 20, 2017

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

drop_duplicates slow for data frame of boolean columns #12963

drop_duplicates slow for data frame of boolean columns #12963

tdhopper commented Apr 22, 2016

sinhrks commented Apr 25, 2016

drop_duplicates slow for data frame of boolean columns #12963

drop_duplicates slow for data frame of boolean columns #12963

Comments

tdhopper commented Apr 22, 2016

sinhrks commented Apr 25, 2016