Subset a data frame based on another

Question

I have two data frames, x and y.

x<-data.frame(id=c(1,2,3,4,5), g=c(21,52,43,94,35))
y<-data.frame(id=c(3,4,7), u=c(55, 77, 99))

I want to subset x to include only the observations with "IDs" that are also in y.

What is the best way of doing this?

Thanks!

Jilber Urbina · Accepted Answer · 2013-09-09T16:34:33.613

6

Use setdiff to exclude observations appearing in both df

> x[setdiff(x$id, y$id),]  
  id  g
1  1 21
2  2 52
5  5 35

Use merge to include observations present in both df

> merge(x, y)
  id  g  u
1  3 43 55
2  4 94 77

or looking for this subset?

> x[intersect(x$id, y$id),]
  id  g
3  3 43
4  4 94

edited Sep 09 '13 at 16:34

answered Sep 09 '13 at 16:27

Jilber Urbina

58,147
10
114
138

1

Your third example was exactly what I was trying to do. Thanks! – Ignacio Sep 09 '13 at 16:41
As the other answerer mentions the `intersect` solution only works by luck. You are subsetting by the ids themselves instead of the indices. – Pierre L Nov 30 '16 at 15:30

score 3 · Answer 2 · answered Nov 29 '16 at 22:20

3

The accepted answer only works because the values 3 and 4 in x$id happen to be located in rows 3 and 4. The wrong answer will be obtained, for example, if:

x<-data.frame(id=c(1,3,2,4,5), g=c(21,52,43,94,35))
x[intersect(x$id, y$id),]
  id  g
3  2 43
4  4 94

The following will work properly, regardless of the position of the common elements:

x[is.element(x$id,intersect(x$id,y$id)),]

answered Nov 29 '16 at 22:20

pbot

141
1
3

1

Nice catch. But the correct fix would be to use the merge solution and avoid `intersect` altogether. Posting a comment under the other answer would suffice. – Pierre L Nov 29 '16 at 22:29
1

The OP wants to subset x. merge(x,y) returns y$u. – pbot Nov 30 '16 at 15:57
Okay I see that now. But that is still better than intersect. `merge(x, y['id'])` – Pierre L Nov 30 '16 at 16:10

Subset a data frame based on another

2 Answers2

Linked