Apache Flume 1.2.x and HBase

The newest (and first) HBase sink was committed into trunk one week ago and was my point at the HBase workshop @Berlin Buzzwords. The slides are available in my slideshare channel.

Let me explain how it works and how you get an Apache Flume - HBase flow running. First, you've got to checkout trunk and build the project (you need git and maven installed on your system):

git clone git://git.apache.org/flume.git && cd flume && git checkout trunk && mvn package -DskipTests && cd flume-ng-dist/target

Within trunk, the HBase sink is available in the sinks - directory (ls -la flume-ng-sinks/flume-ng-hbase-sink/src/main/java/org/apache/flume/sink/hbase/)

Please note a few specialities:
The sink controls atm only HBase flush (), transaction and rollback. Apache Flume reads out the $CLASSPATH variable and uses the first available hbase-site.xml. If you use different versions of HBase on your system please keep that in mind. The HBase table, columns and column family have to be created. Thats all.

The using of an HBase sink is pretty simple, an valid configuration could look like:

host1.sources = src1
host1.sinks = sink1 
host1.channels = ch1 
host1.sources.src1.type = seq 
host1.sources.src1.port = 25001
host1.sources.src1.bind = localhost
host1.sources.src1.channels = ch1
host1.sinks.sink1.type = org.apache.flume.sink.hbase.HBaseSink 
host1.sinks.sink1.channel = ch1
host1.sinks.sink1.table = test3
host1.sinks.sink1.columnFamily = testing
host1.sinks.sink1.column = foo
host1.sinks.sink1.serializer = org.apache.flume.sink.hbase.SimpleHbaseEventSerializer
host1.sinks.sink1.serializer.payloadColumn = pcol
host1.sinks.sink1.serializer.incrementColumn = icol 
host1.channels.ch1.type=memory

In this example we start a Seq interface on localhost with a listening port, point the sink to the HBase sink jar and define the event serializer. Why? HBase needs the data in a HBase format, to achieve that we need to transform the input into a HBase compilant format. Apache Flume's HBase sink uses synchronous / blocking client, asynchronous support will follow (FLUME-1252). 

Links:

Comments

  1. Hi !
    Thanks for informations. I tried that and i have the following result http://pastebin.com/0YZBw8YL . Flume and ZK don't seem to communicate together :(
    Have you an idea ?

    ReplyDelete
  2. Anonymous13 June, 2012

    Hi,

    didn't work on a SASL Zookeeper I guess. Hmm, the certificate are readable? If yes, file pls a jira about.
    http://hbase.apache.org/configuration.html#zk.sasl.auth

    Thanks,
    Alex

    ReplyDelete
  3. Sorry but I don't use SASL ZK. :/

    ReplyDelete
    Replies
    1. Anonymous14 June, 2012

      The pastebin wasn't clear about, so I fired the gun. The nio exception means _mostly_ a heavy used zookeeper cluster. Are HBase and flume installed on the same host? I would point into a problem there. If the error persists please write a mail to the mailinglist.

      Delete
    2. I use HBase on 8 nodes, but it's just for benchmark so my cluster is never used... In addition Zookeeper is deployed on 3 nodes and i have the same problem on all nodes.
      Which mailinglist do you want i notify ?
      Thanks for your work and your help :)

      Delete
    3. That's work now ! Thanks for your helpfull informations.
      Your work is fantasic :)

      Delete
  4. thank you so much alo for the wonderful work..using this writeup I was able to use hbase-sink..but being a newbie I was left some questions..could you please tell something about following 2 line -

    host1.sinks.sink1.serializer.payloadColumn = pcol
    host1.sinks.sink1.serializer.incrementColumn = icol

    Many thanks

    ReplyDelete

Post a Comment

Popular posts from this blog

Export HDFS over CIFS (Samba3)

Hive query shows ERROR "too many counters"

Connect to HiveServer2 with a kerberized JDBC client (Squirrel)