Tuesday, April 28, 2015

Compression In Map Reduce

Compression In Map Reduce
==================

* Compression reduces number of bytes written to/read from HDFS.
* Compression effectively improves the effeciency of network bandwidth and disk space.
* This saves the amount of data being transfored between MAP nodes to REDUCE nodes.

LZO
===

LZO is  a Compression/Decompression Library.

CompressionFomat : LZO
Hadoop CompressionCodec : com.hadoop.compression.lzo.lzopCodec

NOTE: "codec" is the implementation of a compression-decompression alogorithm. In Hadoop, a "codec" is reprasented by an implementation of the "CompressionCodec" interface.

LZO key characterstics:
----------------------------

1. Very fast decompression
2. Requires an additional buffer during the compression (size of 8 kb or 64 kb depends on the compression level)
3. It does not requires the additional buffer during the decompression other than the Source and Destination buffers..thats
    why fast decompresson is possible with LZO.
4. Allows the user to adjust the balance between compression ration and compression speed, without affecting the speed
    of decompression.

--------------------------------------------------------------

Hadoop provides the below Compression Codecs:

- com.hadoop.compression.DefaultCodec
- com.hadoop.compression.lzoCodec
- com.hadoop.compression.SnappyCodec


--------------------------------------------------------------------------------
Configurations Required for LZO compression - In "mapred-site.xml"
--------------------------------------------------------------------------------
By default Compression is not enabled in Mapreduce(i.e. value is false). To achieve the compression we have to edit the below property of 
mapred-site.xml




<property>

      <name>mapred.output.compress</name>
      <value>false</value>

</property>

NOTE: to enable the compression , value should be "true"

----------------------------------------------------------------------

which compression codec to be used while compressing job output. By default "DefaultCodec" will be used.
In order to use other than default(LZO or Snappy) , we have to replace their corresponding codecs.

Like Below:

<property>

      <name>mapred.output.compression.codec</name>
      <value>org.apache.hadoop.io.compress.DefaultCodec</value>

</property>

NOTE: 1.  in place of DefaultCodec, give LzoCodec , SnappyCodec

          2. We can also specify which compression codec to be used while compressing the map outputs.

<property>

      <name>mapred.map.output.compression.codec</name>
      <value>org.apache.hadoop.io.compress.DefaultCodec</value>

</property>  

-------------------------------------------------

Snappy Compression :
==============

   snappy is a also a compression / decompression library. It does not aim for maximum compression, or compatabitlity with
other compression library. 

   Snappy aims for very high speeds and reasonable compression.




Counter In MapReduce

Counters:
======

  Counters are the useful channel for gathering statistics about the job

* for the quality control OR application level statitics
* For Problem diagnosis.

Built-in Counters:
---------------------

Hadoop maintains some built-in counters for every job, which report various metrics  for our job.

Ex:  expected amount of INPUT consumed

      expected amount of OUTPUT produced

Some built in counters:

 1. Map input records --> num of input records consumed by all the maps in the JOB. Incremented every time a record is
                                    read from InputSplit (thru RecordReader) before passing to map() method of Mapper.

2. Map output records--> num of outputt records produced by all the maps in the JOB. Incremented every time a collect()
                                    method is called on Context object

 like wise,

  3. Reduce input records

  4. Reduce output records


** Counters are maintained by the task with which they are associated, and periodically sent to the task tracker and
    then to Job Tracker. So they all can be globally aggregated.

    The built-in Job Counters are actually maintained by the job tracker, so they do not need to be sent across the network
    unlike the all other counters ,including the user defined ones.

Monday, April 27, 2015

USER DEFINED FUNCTION IN PIG HADOOP

Embedded Mode :

 When we are not getting the desired functionality through built in transformation operator in Pig then we ahead with UDF.

Steps For Developing Pig UDF(USER DEFINED FUNCTION)

  1. Write a Class that will extend the base class of EVAL FUNCTION syntax :Eval Function<string>
  2. In order to write business logic we need to overwrite method called execute which takes tuple input 
    exec(tuple input)
  3. We need to dependent compilation Jar File for compilation purpose we have to create our own jar file to be deployed in hadoop environment.
  4. Add the same Jar file Main Script Using REGISTER Keyword.
    Register is the First Line of any script.







Built in transformation or operator in PIG







  • Load -
  • Foreach
  • Generate
  • Filter
  • Dump
  • Store 
  • Split
  • Describe
  • Illustrate
  • Explain
  • Order By
  • Group By
  • Cogroup 
  • Join
  • Union
  • Cross
  • LIMIT
  • Tokenize
  • Flatten
  • AggFunction such as max,min,avg,count etc
  • Distinct


  • Mapreduce VS Apache Pig



    Mapreduce                                                                                     Apache Pig

    Need to write entire logic                                         Built in function are available
    for join ,group,Filter etc

    No of lines of code is too                                           10 lines of pig latin =200
    much for simple fuctionality                                      lines of java


    Time of effort in coding is                                         what took 4 hours to write in
    high.                                                                           it takes 15 min in pig

    Less Productive                                                          High Productive


    next chapter : Built in transformation