find the minimum and maximum date from the data in a RDD in PySpark

Question

I am using Spark with Ipython and have a RDD which contains data in this format when printed:

print rdd1.collect()

[u'2010-12-08 00:00:00', u'2010-12-18 01:20:00', u'2012-05-13 00:00:00',....]

Each data is a datetimestamp and I want to find the minimum and the maximum in this RDD. How can I do that?

score 7 · Answer 1 · edited May 23 '17 at 12:19

You can for example use aggregate function (for an explanation how it works see: What is the equivalent implementation of RDD.groupByKey() using RDD.aggregateByKey()?)

from datetime import datetime    

rdd  = sc.parallelize([
    u'2010-12-08 00:00:00', u'2010-12-18 01:20:00', u'2012-05-13 00:00:00'])

def seq_op(acc, x):
    """ Given a tuple (min-so-far, max-so-far) and a date string
    return a tuple (min-including-current, max-including-current)
    """
    d = datetime.strptime(x, '%Y-%m-%d %H:%M:%S')
    return (min(d, acc[0]), max(d, acc[1]))

def comb_op(acc1, acc2):
    """ Given a pair of tuples (min-so-far, max-so-far)
    return a tuple (min-of-mins, max-of-maxs)
    """
    return (min(acc1[0], acc2[0]), max(acc1[1], acc2[1]))

# (initial-min <- max-date, initial-max <- min-date)
rdd.aggregate((datetime.max, datetime.min), seq_op, comb_op)

## (datetime.datetime(2010, 12, 8, 0, 0), datetime.datetime(2012, 5, 13, 0, 0))

or DataFrames:

from pyspark.sql import Row
from pyspark.sql.functions import from_unixtime, unix_timestamp, min, max

row = Row("ts")
df = rdd.map(row).toDF()

df.withColumn("ts", unix_timestamp("ts")).agg(
    from_unixtime(min("ts")).alias("min_ts"), 
    from_unixtime(max("ts")).alias("max_ts")
).show()

## +-------------------+-------------------+
## |             min_ts|             max_ts|
## +-------------------+-------------------+
## |2010-12-08 00:00:00|2012-05-13 00:00:00|
## +-------------------+-------------------+

is it possible if you could mention some comments about how is the code working (the two functions especially) so that I am able to understand it properly? — Jason Donnald, Nov 25 '15 at 04:57

lanenok · Answer 2 · 2015-11-25T10:32:10.127

3

If you RDD consists of datetime objects, what is wrong with simply using

rdd1.min()
rdd1.max()

See documentation

This example works for me

rdd = sc.parallelize([u'2010-12-08 00:00:00', u'2010-12-18 01:20:00', u'2012-05-13 00:00:00'])
from datetime import datetime
rddT = rdd.map(lambda x: datetime.strptime(x, "%Y-%m-%d %H:%M:%S")).cache()
print rddT.min()
print rddT.max()

edited Nov 25 '15 at 10:32

answered Nov 25 '15 at 08:50

lanenok

2,699
17
24

score 1 · Answer 3 · answered Sep 17 '19 at 14:07

1

If you're using Dataframes you don't need nothing but this:

import pyspark.sql.functions as F 
#Anohter imports, session, attributions etc
# This brings min and max(considering you need only min and max values)
df.select(F.min('datetime_column_name'),F.max('datetime_column_name')).show()

and that's it!

answered Sep 17 '19 at 14:07

Andre Carneiro

708
1
5
27

It is simple and worked for me! Thank you – daz Jul 06 '21 at 15:21

find the minimum and maximum date from the data in a RDD in PySpark

3 Answers3

Linked