When to use SPARK_CLASSPATH or SparkContext.addJar

Question

I'm using a standalone spark cluster, one master and 2 workers. I really don't understand how to use wisely SPARK_CLASSPATH or SparkContext.addJar. I tried both and It looks like addJar doesn't work as I used to believe.

In my case I tried to use some joda-time function, in the closures or outside. If I set SPARK_CLASSPATH with a path to the joda-time jar, everything works ok. But if I remove SPARK_CLASSPATH and add in my program:

JavaSparkContext sc = new JavaSparkContext("spark://localhost:7077", "name", "path-to-spark-home", "path-to-the-job-jar");
sc.addJar("path-to-joda-jar");

It doesn't work anymore, although in logs I can see:

14/03/17 15:32:57 INFO SparkContext: Added JAR /home/hduser/projects/joda-time-2.1.jar at http://127.0.0.1:46388/jars/joda-time-2.1.jar with timestamp 1395066777041

and immediatly after:

Caused by: java.lang.NoClassDefFoundError: org/joda/time/DateTime
    at com.xxx.sparkjava1.SimpleApp.main(SimpleApp.java:57)
    ... 6 more
Caused by: java.lang.ClassNotFoundException: org.joda.time.DateTime
    at java.net.URLClassLoader$1.run(URLClassLoader.java:366)

I used to suppose that SPARK_CLASSPATH was setting the classpath for the driver part of the job, and SparkContext.addJar was setting the classpath for the executors, but It does not seem right anymore.

Anyone knows better than me?

`SPARK_CLASSPATH`is deprecated since Spark 1.0+. http://stackoverflow.com/questions/37132559/add-jars-to-a-spark-job-spark-submit — Daniel Carroza, Nov 08 '16 at 12:58
Try this already answered post. this provides good description. https://stackoverflow.com/questions/37132559/add-jars-to-a-spark-job-spark-submit — VimalK, Mar 04 '21 at 07:22

score 1 · Accepted Answer · answered Mar 18 '14 at 16:58

SparkContext.addJar is broken in 0.9 as well as ADD_JARS environment variable. It used to work as documented in 0.8.x and the fix is already commited to master, so it's expected in the next release. For now you can either use workaround described in Jira or make patched Spark build.

See relevant mailing list discussion: http://mail-archives.apache.org/mod_mbox/spark-user/201402.mbox/%3C5234E529519F4320A322B80FBCF5BDA6@gmail.com%3E

Jira issue: https://spark-project.atlassian.net/plugins/servlet/mobile#issue/SPARK-1089

Ok, so it's a bug and the workaround is to use SPARK_CLASSPATH. Thanks. — VirgileD, Mar 19 '14 at 08:32

score 0 · Answer 2 · edited May 23 '17 at 12:06

0

SPARK_CLASSPATH is deprecated since Spark 1.0+. You can add jars to the classpath programatically, inside file spark-defaults.conf or with spark-submit flags.

Add jars to a Spark Job - spark-submit

edited May 23 '17 at 12:06

Community

1
1

answered Nov 08 '16 at 13:09

Daniel Carroza

310
1
8

When to use SPARK_CLASSPATH or SparkContext.addJar

2 Answers2