Thursday, May 16, 2013

Play run in DEV mode and "ClassNotFoundException"

Play "~run" command makes development much easier. You can change the code and test it without building, packaging and deploying. But it also causes annoying "ClassNotFoundException".

Our application has a hbase filter. The filter will be packaged and deployed to HBase region servers, but it is also needed on the client side when you build a Scan. If we run the application in DEV mode, we will get "ClassNotFoundException". The java code of the filter is definitely compiled and "in the classpath" because I can find it in the output of "show full-classpath". This confusing issue forces us to use stage/dist again.

The issue is actually caused by the classloader when you start "run" command. If you use a customized filter, HBase will use "Class.forName" to load the class. Because the filter is NOT in the classpath of the classloader which loads HBase classes, "ClassNotFoundException" is thrown.

But Why the filter is NOT in the classpath? There are several classloaders when Play runs in DEV mode:

  • sbtLoader, the loader loads sbt;
  • applicationLoader, the loader loads the jar files in dependencyClasspath in Compile. it is also called as "SBT/Play shared ClassLoader".
  • ReloadableClassLoader(v1), the loader loads the classes of the project
Because ReloadableClassLoader is a child of applicationLoader, the filter is in the classpath of ReloadableClassLoader rather than applicationLoader, and HBase library uses applicationLoader, the filter is indeed invisible to applicationLoader.

The simple workaround is to make the filter a separate project and a dependency of play. The only disadvantage is you have to build, publish, and update if you are developing the filter and the application code at the same time.

You can find the same issue: https://github.com/playframework/Play20/issues/822

Play 2.1.1 Jdbc is not compatible with hive-jdbc-0.10.0-cdh4.2.0

If you want to access a Hive/Impala server in Play 2.1.1, you will encounter a strange error saying "Cannot connect to database [xxx]. Unfortunately no further information can help you identify where is wrong. It is strange because you can using the same URL to connect to the server without any problem.
import java.sql.DriverManager

object ImpalaJdbcClient extends App {

  val impalaUrl = "jdbc:hive2://impala-host:21050/;auth=noSasl"

  val driverClass = Class.forName("org.apache.hive.jdbc.HiveDriver")

  val conn = DriverManager.getConnection(impalaUrl)

  val st = conn.createStatement()
  val rs = st.executeQuery("select count(distinct customer_id) from customers where repeat ='Y'")
  while (rs.next()) {
    println("count=%d".format(rs.getLong(1)))
  }

  conn.close
}
By digging Play and Hive-JDBC code, I figured out that Play-jdbc calls a lot of methods what hive-jdbc doesn't support. For those methods, such as setReadOnly and setCatalog, hive-jdbc just simply throws a SQLException saying "Method not supported", then Play-jdbc catch it and report "Cannot connect to database" error, but unfortunately it doesn't include the message of "Method not supported". You can fix it by removing throw statements from hive-jdbc unsupported method and recompiling. Another way is to create your own BoneCPPlugin. Just copy the source code ./src/play-jdbc/src/main/scala/play/api/db/DB.scala and remove the offending method calls:
  • setAutoCommit
  • commit
  • rollback
  • setTransactionIsolation
  • setReadOnly
  • setCatalog
and comment or replace this line
case mode => Logger("play").info("database [" + ds._2 + "] connected at " + dbURL(ds._1.getConnection))
to
case mode => Logger("play").info("database [" + ds._2 + "] connected at " + ds._1)
because dbURL calls conn.getMetaData.getURL and HiveDatabaseMetaData doesn't support getURL. Change dbplugin in app.configuration.getString("dbplugin").filter(_ == "disabled") to something else to avoid conflict with Play's BoneCPPlugin. Then register your own BoneCPPlugin in conf/play.plugins.

Wednesday, May 15, 2013

How Play set up ivy repository?

Play sets ivy repository to ${PLAY_HOME}/repository instead of using the default Ivy home $HOME/.ivy2. Checkout the file ${PLAY_HOME}/play and you will find this at the end
"$JAVA" -Dsbt.ivy.home=$dir/repository -Dplay.home=$dir/framework -Dsbt.boot.properties=$dir/framework/sbt/play.boot.properties -jar $dir/framework/sbt/sbt-launch.jar "$@"
-Dsbt.ivy.home does the trick.

Wednesday, May 1, 2013

Install Impala 1.0 in Cloudera Manager 4.5.0

If you don't want to upgrade to 4.5.2, you can change impala parcel URL to get the impala parcel. http://archive.cloudera.com/impala/parcels/latest/

Monday, April 29, 2013

Run MapReduce in Play development mode

When you want to invoke a MapReduce job in Play development mode, the first problem you have will be that the jar file is not generated yet. so job.setJarByClass(classOf[Mapper]) won't work. You can overcome this issue by using Hadoop utility class JarFinder like this:
    job.setJarByClass(classOf[MyMapper])
    if (job.getJar() == null) {
      val jarFinder = Class.forName("org.apache.hadoop.util.JarFinder")
      if (jarFinder != null) {
        val m = jarFinder.getMethod("getJar", classOf[Class[_]]);

        job.getConfiguration().asInstanceOf[org.apache.hadoop.mapred.JobConf]
          .setJar(m.invoke(null, classOf[MyMapper]).asInstanceOf[String])
      }
    }
Here are what you need to pay attentions:
  • org.apache.hadoop.util.JarFinder is a Hadoop Utility class for testing, which creates a Jar on the fly. You can find the jar in hadoop-common package with classifier "tests", hadoop-common-2.0.0-cdh4.2.1-tests.jar in cloudera distribution. Unfortunately simply putting a dependency into Build.scala won't work:
       "org.apache.hadoop" % "hadoop-common" % "2.0.0-cdh4.2.1" classifier "tests", 
    
    You will get a lot of compilation errors. If you add "test" like this, compilation pass but you cannot use JarFinder in development mode because the jar is not in the classpath (it IS in test classpath)
       "org.apache.hadoop" % "hadoop-common" % "2.0.0-cdh4.2.1" % "test" classifier "tests", 
    
    I just simply get JarFinder.java from the source package and put into play app/org/apache/hadoop/util directory like other scala files. It will be compiled by play.
  • Depends on which hadoop version you are using, you may use setJar directly on job like this
      job.setJar(m.invoke(null, classOf[MyMapper]).asInstanceOf[String])
    
    I'm using "org.apache.hadoop" % "hadoop-client" % "2.0.0-mr1-cdh4.2.0". org.apache.hadoop.mapreduce.Job doesn't have setJar.
  • Property mapred.jar will be created when setJar is called, you can verify if this property exists in the job file on JobTracker web page. You could make a mistake to setJar in another Configuration object which is not the one used by job, for example,
        val config = HBaseConfiguration.create()
        val job = new Job(config, "My MR job")
        ...
        job.setJarByClass(classOf[MyMapper])
        if (job.getJar() == null) {
          val jarFinder = Class.forName("org.apache.hadoop.util.JarFinder")
          if (jarFinder != null) {
            ...
            // This won't work because jc is a cloned object, not used by Job.
            // You cannot use val config too for the same reason.
            val jc = new JobConf(job.getConfiguration())
            jc.setJar(m.invoke(null, classOf[MyMapper]).asInstanceOf[String]);
          }
        }
    
  • Another problem you may encounter is that JarFinder.getJar still returns null. I had this problem when I ran a HBase sbt project, but don't remember if this happened in Play project. If you have this problem, you can add the following code in JarFinder to fix it:
      public static String getJar(Class klass) {
        Preconditions.checkNotNull(klass, "klass");
        ClassLoader loader = klass.getClassLoader();
    
        // Try to use context class loader 
        if (loader == null) {
         loader = Thread.currentThread().getContextClassLoader();
        }
    
        if (loader != null) {
    
    My HBase project starts a MR job that needs scala-library.jar distributed. Here is the snippet
    
        TableMapReduceUtil.initTableMapperJob(
          tableName,
          getScan(siteKey, collDefId),
          classOf[TextMapper],
          classOf[ImmutableBytesWritable],
          classOf[Text],
          job,
          // don't add dependency jars
          false)
        job.setOutputFormatClass(classOf[SequenceFileOutputFormat[Text, Text]])
        job.setOutputKeyClass(classOf[Text])
        job.setOutputValueClass(classOf[Text])
        FileOutputFormat.setCompressOutput(job, true)
        FileOutputFormat.setOutputPath(job, new Path(outDir))
        job.setNumReduceTasks(0);
    
        // Add dependency jars 
        TableMapReduceUtil.addDependencyJars(job.getConfiguration(),
          classOf[org.apache.zookeeper.ZooKeeper],
          classOf[com.google.protobuf.Message],
          job.getMapOutputKeyClass(),
          job.getMapOutputValueClass(),
          job.getInputFormatClass(),
          job.getOutputKeyClass(),
          job.getOutputValueClass(),
          job.getOutputFormatClass(),
          job.getPartitionerClass(),
          job.getCombinerClass(),
    
          // include scala-library.jar
          classOf[ScalaObject],
    
          // joda jar used by mapper/reducer
          classOf[org.joda.time.DateTime]);
    
    HBase TableMapReduceUtil (0.94.2-cdh4.2.0) will use JarFinder.getJar if it presents. Unfortunately the classloader used by sbt to launcher the application doesn't have ScalaObject in the classpath. klass.getClassLoader() will return null. When you use the context classloader, scala-library will be found without any problem.
  • The code works in both dev and prod mode. In prod mode, the jar of mappers and reducers are already in the classpath, setJarByClass will work without entering JarFinder block.

Monday, April 1, 2013

Extract compressed Tar

  • Extract GZIP compressed TAR file
    tar xzvf file.tar.gz
  • Extract BZ2 compressed TAR file
    tar xvjpf file.tar.gz

Friday, March 1, 2013

Git-SVN

Here are some notes when I play with git-svn:
  • Clone a project. You can run the following commands if the project follows the standard layout with trunk/tags/branches.
    root
      |- ProjectA
      |   |- branches
      |   |- tags
      |   \- trunk
      |- OtherProject
    
  • Download trunk, tags and branches. You will find the same directories just like SVN repository.
    git svn clone -r XXXX https://subversion.mycompany.com/ProjectA
    
  • init a new project. This will create dir projectA and initializ it as a git repository.
    git svn init https://subversion.mycompany.com/projectA projectA
    
  • Download trunk only. The directory have the files under trunk only.
    git svn clone -s -r XXXX https://subversion.mycompany.com/ProjectA
    • Provide the revision number using "-r xxxx", otherwise have to wait for a long time until git-svn finishes scanning all revisions.
    • Use the revision of the project (96365) instead of the latest revision (96381). Otherwise, you will get the working directory empty. You can find the revision number in viewvc like this:
      Index of /ProjectA
      
      Files shown: 0
      Directory revision: 96365 (of 96381)
      Sticky Revision:    
      
  • Apply git patch to SVN:
    git diff --no-prefix > /tmp/xx.patch
    cd svn-dir/
    patch -p0 < /tmp/xx.patch