Are there some ready to use libraries or packages for python or R to reduce the number of levels for large categorical factors?
I want to achieve something similar to R: "Binning" categorical variables but encode into the most frequently top-k factors and "other".