Wednesday, May 1, 2013
Install Impala 1.0 in Cloudera Manager 4.5.0
If you don't want to upgrade to 4.5.2, you can change impala parcel URL to get the impala parcel.
http://archive.cloudera.com/impala/parcels/latest/
Monday, April 29, 2013
Run MapReduce in Play development mode
When you want to invoke a MapReduce job in Play development mode, the first problem you have will be that the jar file is not generated yet. so job.setJarByClass(classOf[Mapper]) won't work. You can overcome this issue by using Hadoop utility class JarFinder like this:
job.setJarByClass(classOf[MyMapper])
if (job.getJar() == null) {
val jarFinder = Class.forName("org.apache.hadoop.util.JarFinder")
if (jarFinder != null) {
val m = jarFinder.getMethod("getJar", classOf[Class[_]]);
job.getConfiguration().asInstanceOf[org.apache.hadoop.mapred.JobConf]
.setJar(m.invoke(null, classOf[MyMapper]).asInstanceOf[String])
}
}
Here are what you need to pay attentions:
- org.apache.hadoop.util.JarFinder is a Hadoop Utility class for testing, which creates a Jar on the fly. You can find the jar in hadoop-common package with classifier "tests", hadoop-common-2.0.0-cdh4.2.1-tests.jar in cloudera distribution. Unfortunately simply putting a dependency into Build.scala won't work:
"org.apache.hadoop" % "hadoop-common" % "2.0.0-cdh4.2.1" classifier "tests",
You will get a lot of compilation errors. If you add "test" like this, compilation pass but you cannot use JarFinder in development mode because the jar is not in the classpath (it IS in test classpath)"org.apache.hadoop" % "hadoop-common" % "2.0.0-cdh4.2.1" % "test" classifier "tests",
I just simply get JarFinder.java from the source package and put into play app/org/apache/hadoop/util directory like other scala files. It will be compiled by play. - Depends on which hadoop version you are using, you may use setJar directly on job like this
job.setJar(m.invoke(null, classOf[MyMapper]).asInstanceOf[String])
I'm using "org.apache.hadoop" % "hadoop-client" % "2.0.0-mr1-cdh4.2.0". org.apache.hadoop.mapreduce.Job doesn't have setJar. - Property mapred.jar will be created when setJar is called, you can verify if this property exists in the job file on JobTracker web page. You could make a mistake to setJar in another Configuration object which is not the one used by job, for example,
val config = HBaseConfiguration.create() val job = new Job(config, "My MR job") ... job.setJarByClass(classOf[MyMapper]) if (job.getJar() == null) { val jarFinder = Class.forName("org.apache.hadoop.util.JarFinder") if (jarFinder != null) { ... // This won't work because jc is a cloned object, not used by Job. // You cannot use val config too for the same reason. val jc = new JobConf(job.getConfiguration()) jc.setJar(m.invoke(null, classOf[MyMapper]).asInstanceOf[String]); } } - Another problem you may encounter is that JarFinder.getJar still returns null. I had this problem when I ran a HBase sbt project, but don't remember if this happened in Play project. If you have this problem, you can add the following code in JarFinder to fix it:
public static String getJar(Class klass) { Preconditions.checkNotNull(klass, "klass"); ClassLoader loader = klass.getClassLoader(); // Try to use context class loader if (loader == null) { loader = Thread.currentThread().getContextClassLoader(); } if (loader != null) {My HBase project starts a MR job that needs scala-library.jar distributed. Here is the snippetTableMapReduceUtil.initTableMapperJob( tableName, getScan(siteKey, collDefId), classOf[TextMapper], classOf[ImmutableBytesWritable], classOf[Text], job, // don't add dependency jars false) job.setOutputFormatClass(classOf[SequenceFileOutputFormat[Text, Text]]) job.setOutputKeyClass(classOf[Text]) job.setOutputValueClass(classOf[Text]) FileOutputFormat.setCompressOutput(job, true) FileOutputFormat.setOutputPath(job, new Path(outDir)) job.setNumReduceTasks(0); // Add dependency jars TableMapReduceUtil.addDependencyJars(job.getConfiguration(), classOf[org.apache.zookeeper.ZooKeeper], classOf[com.google.protobuf.Message], job.getMapOutputKeyClass(), job.getMapOutputValueClass(), job.getInputFormatClass(), job.getOutputKeyClass(), job.getOutputValueClass(), job.getOutputFormatClass(), job.getPartitionerClass(), job.getCombinerClass(), // include scala-library.jar classOf[ScalaObject], // joda jar used by mapper/reducer classOf[org.joda.time.DateTime]);HBase TableMapReduceUtil (0.94.2-cdh4.2.0) will use JarFinder.getJar if it presents. Unfortunately the classloader used by sbt to launcher the application doesn't have ScalaObject in the classpath. klass.getClassLoader() will return null. When you use the context classloader, scala-library will be found without any problem. - The code works in both dev and prod mode. In prod mode, the jar of mappers and reducers are already in the classpath, setJarByClass will work without entering JarFinder block.
Monday, April 1, 2013
Extract compressed Tar
- Extract GZIP compressed TAR file
tar xzvf file.tar.gz
- Extract BZ2 compressed TAR file
tar xvjpf file.tar.gz
Friday, March 1, 2013
Git-SVN
Here are some notes when I play with git-svn:
- Clone a project. You can run the following commands if the project follows the standard layout with trunk/tags/branches.
root |- ProjectA | |- branches | |- tags | \- trunk |- OtherProject
- Download trunk, tags and branches. You will find the same directories just like SVN repository.
git svn clone -r XXXX https://subversion.mycompany.com/ProjectA
- init a new project. This will create dir projectA and initializ it as a git repository.
git svn init https://subversion.mycompany.com/projectA projectA
- Download trunk only. The directory have the files under trunk only.
git svn clone -s -r XXXX https://subversion.mycompany.com/ProjectA
- Provide the revision number using "-r xxxx", otherwise have to wait for a long time until git-svn finishes scanning all revisions.
- Use the revision of the project (96365) instead of the latest revision (96381). Otherwise, you will get the working directory empty. You can find the revision number in viewvc like this:
Index of /ProjectA Files shown: 0 Directory revision: 96365 (of 96381) Sticky Revision:
- Apply git patch to SVN:
git diff --no-prefix > /tmp/xx.patch cd svn-dir/ patch -p0 < /tmp/xx.patch
Wednesday, February 20, 2013
Friday, February 15, 2013
Scala SBT
Requirements:
- Include scala-library.jar in the zip file because the machine may not have scala installed.
- Include scripts and configure files into the zip just like Maven assembly.
- Include the jar in the zip file.
Lessons:
- Key.
- SettingKey and TaskKey
- Cannot reference Key directory. Need to use tuple like this:
(sbt_key) => { key_val => } - projectBin is the jar
- managed
- Must have a Project
- libraryDependencies and resolvers defined in Build not effect for the project
- IO.zip doesn't support permission of shell scripts
- Build with java code perfectly.
- custom cleanFiles
- Define a custom task like
dist - Useful commands of sbt
settings // list settings keys tasks // list all tasks show clean-files // check value of settings keys inspect clean-files // check more information than show
HBase LeaseException issue
Excerpt from http://hbase.apache.org/book.html. If you google "hbase LeaseException", this page may not be on the first page.
12.5.2.
12.5.2. LeaseException when calling Scanner.next
In some situations clients that fetch data from a RegionServer get a LeaseException instead of the usual Section 12.5.1, “ScannerTimeoutException or UnknownScannerException”. Usually the source of the exception is
org.apache.hadoop.hbase.regionserver.Leases.removeLease(Leases.java:230) (line number may vary). It tends to happen in the context of a slow/freezing RegionServer#next call. It can be prevented by having hbase.rpc.timeout > hbase.regionserver.lease.period. Harsh J investigated the issue as part of the mailing list thread HBase, mail # user - Lease does not exist exceptions
Subscribe to:
Posts (Atom)