Compression In Map Reduce
==================
* Compression reduces number of bytes written to/read from HDFS.
* Compression effectively improves the effeciency of network bandwidth and disk space.
* This saves the amount of data being transfored between MAP nodes to REDUCE nodes.
LZO
===
LZO is a Compression/Decompression Library.
CompressionFomat : LZO
Hadoop CompressionCodec : com.hadoop.compression.lzo.lzopCodec
NOTE: "codec" is the implementation of a compression-decompression alogorithm. In Hadoop, a "codec" is reprasented by an implementation of the "CompressionCodec" interface.
LZO key characterstics:
----------------------------
1. Very fast decompression
2. Requires an additional buffer during the compression (size of 8 kb or 64 kb depends on the compression level)
3. It does not requires the additional buffer during the decompression other than the Source and Destination buffers..thats
why fast decompresson is possible with LZO.
4. Allows the user to adjust the balance between compression ration and compression speed, without affecting the speed
of decompression.
--------------------------------------------------------------
Hadoop provides the below Compression Codecs:
- com.hadoop.compression.DefaultCodec
- com.hadoop.compression.lzoCodec
- com.hadoop.compression.SnappyCodec
--------------------------------------------------------------------------------
Configurations Required for LZO compression - In "mapred-site.xml"
--------------------------------------------------------------------------------
By default Compression is not enabled in Mapreduce(i.e. value is false). To achieve the compression we have to edit the below property of
mapred-site.xml
<property>
<name>mapred.output.compress</name>
<value>false</value>
</property>
NOTE: to enable the compression , value should be "true"
----------------------------------------------------------------------
which compression codec to be used while compressing job output. By default "DefaultCodec" will be used.
In order to use other than default(LZO or Snappy) , we have to replace their corresponding codecs.
Like Below:
<property>
<name>mapred.output.compression.codec</name>
<value>org.apache.hadoop.io.compress.DefaultCodec</value>
</property>
NOTE: 1. in place of DefaultCodec, give LzoCodec , SnappyCodec
2. We can also specify which compression codec to be used while compressing the map outputs.
<property>
<name>mapred.map.output.compression.codec</name>
<value>org.apache.hadoop.io.compress.DefaultCodec</value>
</property>
-------------------------------------------------
Snappy Compression :
==============
snappy is a also a compression / decompression library. It does not aim for maximum compression, or compatabitlity with
other compression library.
Snappy aims for very high speeds and reasonable compression.
ASP.NET is a web application framework developed and marketed by Microsoft to allow programmers to build dynamic web sites, web applications and web services. It was first released in January 2002 with version 1.0 of the .NET Framework, and is the successor to Microsoft's Active Server Pages (ASP) technology. ASP.NET is built on the Common Language Runtime (CLR), allowing programmers to write ASP.NET code using any supported .NET language.
Tuesday, April 28, 2015
Counter In MapReduce
Counters:
======
Counters are the useful channel for gathering statistics about the job
* for the quality control OR application level statitics
* For Problem diagnosis.
Built-in Counters:
---------------------
Hadoop maintains some built-in counters for every job, which report various metrics for our job.
Ex: expected amount of INPUT consumed
expected amount of OUTPUT produced
Some built in counters:
1. Map input records --> num of input records consumed by all the maps in the JOB. Incremented every time a record is
read from InputSplit (thru RecordReader) before passing to map() method of Mapper.
2. Map output records--> num of outputt records produced by all the maps in the JOB. Incremented every time a collect()
method is called on Context object
like wise,
3. Reduce input records
4. Reduce output records
** Counters are maintained by the task with which they are associated, and periodically sent to the task tracker and
then to Job Tracker. So they all can be globally aggregated.
The built-in Job Counters are actually maintained by the job tracker, so they do not need to be sent across the network
unlike the all other counters ,including the user defined ones.
======
Counters are the useful channel for gathering statistics about the job
* for the quality control OR application level statitics
* For Problem diagnosis.
Built-in Counters:
---------------------
Hadoop maintains some built-in counters for every job, which report various metrics for our job.
Ex: expected amount of INPUT consumed
expected amount of OUTPUT produced
Some built in counters:
1. Map input records --> num of input records consumed by all the maps in the JOB. Incremented every time a record is
read from InputSplit (thru RecordReader) before passing to map() method of Mapper.
2. Map output records--> num of outputt records produced by all the maps in the JOB. Incremented every time a collect()
method is called on Context object
like wise,
3. Reduce input records
4. Reduce output records
** Counters are maintained by the task with which they are associated, and periodically sent to the task tracker and
then to Job Tracker. So they all can be globally aggregated.
The built-in Job Counters are actually maintained by the job tracker, so they do not need to be sent across the network
unlike the all other counters ,including the user defined ones.
Monday, April 27, 2015
USER DEFINED FUNCTION IN PIG HADOOP
Embedded Mode :
When we are not getting the desired functionality through built in transformation operator in Pig then we ahead with UDF.
Steps For Developing Pig UDF(USER DEFINED FUNCTION)
When we are not getting the desired functionality through built in transformation operator in Pig then we ahead with UDF.
Steps For Developing Pig UDF(USER DEFINED FUNCTION)
- Write a Class that will extend the base class of EVAL FUNCTION syntax :Eval Function<string>
- In order to write business logic we need to overwrite method called execute which takes tuple input
exec(tuple input) - We need to dependent compilation Jar File for compilation purpose we have to create our own jar file to be deployed in hadoop environment.
- Add the same Jar file Main Script Using REGISTER Keyword.
Register is the First Line of any script.
Built in transformation or operator in PIG
Next Chapter : User Defind Function in Apache Pig
Mapreduce VS Apache Pig
Mapreduce Apache Pig
Need to write entire logic Built in function are available
for join ,group,Filter etc
No of lines of code is too 10 lines of pig latin =200
much for simple fuctionality lines of java
Time of effort in coding is what took 4 hours to write in
high. it takes 15 min in pig
Less Productive High Productive
next chapter : Built in transformation
Subscribe to:
Posts (Atom)