Pandas groupby - set of different values

Question

I have this dataframe

x = pd.DataFrame.from_dict({'cat1':['A', 'A', 'A', 'B', 'B', 'C', 'C', 'C'], 'cat2':['X', 'X', 'Y', 'Y', 'Y', 'Y', 'Z', 'Z']})

  cat1 cat2
0    A    X
1    A    X
2    A    Y
3    B    Y
4    B    Y
5    C    Y
6    C    Z
7    C    Z

I want to group by cat1, and then aggregate cat2 as sets of different values, such as

  cat1 cat2
0    A    (X, Y)
1    B    (Y,)
2    C    (Y, Z)

This is part of a bigger dataframe with more columns, each of which has its own aggregation function, so how do I pass this functionality to the aggregation dictionary?

jezrael · Accepted Answer · 2017-11-29T14:13:52.473

Use lambda function with set or unique, also convert output to tuples:

x = pd.DataFrame.from_dict({'cat1':['A', 'A', 'A', 'B', 'B', 'C', 'C', 'C'], 
                            'cat2':['X', 'X', 'Y', 'Y', 'Y', 'Y', 'Z', 'Z'],
                             'col':range(8)})
print (x)
  cat1 cat2  col
0    A    X    0
1    A    X    1
2    A    Y    2
3    B    Y    3
4    B    Y    4
5    C    Y    5
6    C    Z    6
7    C    Z    7

a = x.groupby('cat1').agg({'cat2': lambda x: tuple(set(x)), 'col':'sum'})
print (a)
        cat2  col
cat1             
A     (Y, X)    3
B       (Y,)    7
C     (Y, Z)   18

Or:

a = x.groupby('cat1').agg({'cat2': lambda x: tuple(x.unique()), 'col':'sum'})
print (a)
        cat2  col
cat1             
A     (X, Y)    3
B       (Y,)    7
C     (Y, Z)   18

EDIT:

f = lambda x: tuple(x.unique())
f.__name__ = 'my_name'
a = x.groupby('cat1')['cat2'].agg(['min', 'max', 'nunique', f])
print (a)
     min max  nunique my_name
cat1                         
A      X   Y        2  (X, Y)
B      Y   Y        1    (Y,)
C      Y   Z        2  (Y, Z)

If there is only one lambda function or no problem with column name <lambda>:

a = x.groupby('cat1')['cat2'].agg(['min', 'max', 'nunique', lambda x: tuple(x.unique())])
print (a)
     min max  nunique <lambda>
cat1                          
A      X   Y        2   (X, Y)
B      Y   Y        1     (Y,)
C      Y   Z        2   (Y, Z)

So I already have something like 'cat2': ['min', 'max', 'nunique'], i.e. I'm already aggregating this column multiple ways. How can your solution be modified to accomodate for this? Thanks — Baron Yugovich, Nov 29 '17 at 14:04
There is possible use custom function and set name by `__name__`, check last edit. — jezrael, Nov 29 '17 at 14:07

score 3 · Answer 2 · answered Nov 29 '17 at 00:18

3

Groupby and unique gives you unique values

x.groupby('cat1').cat2.unique()

A    [X, Y]
B       [Y]
C    [Y, Z]

If you want to have the output in tuple, try

x.groupby('cat1').cat2.unique().apply(tuple)

A    (X, Y)
B      (Y,)
C    (Y, Z)

answered Nov 29 '17 at 00:18

Vaishali

37,545
5
58
86

Please see my edits to the question above. I need to do this as part of a bigger aggregation dictionary. – Baron Yugovich Nov 29 '17 at 13:47

score 3 · Answer 3 · answered Nov 29 '17 at 00:20

3

x.groupby('cat1')['cat2'].unique().reset_index()

# Returns 
  cat1    cat2
0    A  [X, Y]
1    B     [Y]
2    C  [Y, Z]

This first groups the entire dataframe by 'cat1', selects only the series 'cat2', and reduces each group to the unique set of 'cat2' values. The result puts the 'cat1' values in the index, so reset_index() will pull those values back out as a column if you need it in that format.

answered Nov 29 '17 at 00:20

Simon Bowly

1,003
5
10

Please see my edits to the question above. I need to do this as part of a bigger aggregation dictionary. – Baron Yugovich Nov 29 '17 at 13:47

score 2 · Answer 4 · edited Jan 04 '21 at 16:50

2

x.groupby('cat1')['cat2'].agg(lambda x: set(x))

Output

As for the simplification suggested in comments, it appears the following works at least with Python 3.6.5 and Pandas 0.23.0 (but not with Python 3.6.2 and Pandas 0.20.3) :

x.groupby('cat1')['cat2'].agg(set)

edited Jan 04 '21 at 16:50

Skippy le Grand Gourou

6,976
4
60
76

answered Nov 29 '17 at 00:17

Alter

3,332
4
31
56

1

The lambda is not really necessary here, since `set` is callable. So x.groupby('cat1').agg(set) does the same thing, doesn't it? – Simon Bowly Nov 29 '17 at 00:23
1

It doesn't work in this case, though I thought it would too – Alter Nov 29 '17 at 00:26
Please see my edits to the question above. I need to do this as part of a bigger aggregation dictionary. – Baron Yugovich Nov 29 '17 at 13:47
1

This solution is WAY faster than solutions based on `unique` and `apply`. – Skippy le Grand Gourou Jan 04 '21 at 16:10

score 2 · Answer 5 · answered Nov 29 '17 at 02:09

2

Or we can filter the dataframe before groupby

x.drop_duplicates().groupby('cat1').cat2.apply(tuple)
Out[777]: 
cat1
A    (X, Y)
B      (Y,)
C    (Y, Z)
Name: cat2, dtype: object

answered Nov 29 '17 at 02:09

BENY

317,841
20
164
234

Please see my edits to the question above. I need to do this as part of a bigger aggregation dictionary. – Baron Yugovich Nov 29 '17 at 13:47

Pandas groupby - set of different values

5 Answers5

Linked