Showing posts with label Interview Questions. Show all posts
Showing posts with label Interview Questions. Show all posts

Friday, 15 May 2015

Hadoop Interview questions 2015 - Part3

1. How do you handle exceptions in MapReduce?
2. Modeling in NoSQL
3. Explain about one of the MapReduce program that you have implemented which u think it can't be done with Hive/Pig
4. What kind of initial validations that you will do on data before processing?
5. Explain when to choose MapReduce?
6. Differences b/w MapR 1 & 2
7. What is HA node in hadoop?  Do we have this in MapR 1?
8. What are the issues you faced in hadoop projects?
9.Difference between MapR and Cloudera distributions
10. What are the factors considered to decide key and the key format in mapreduce?

Saturday, 25 April 2015

Pig Interview Questions 2015

Pig Interview questions

1. How pig calculates the number of reducers for a job
2. How to change the default reducer calculation method in pig
3. What is hash based aggregation
4. Performance optimizations in pig
5. Types of joins in pig
6. Explain about merge-sparse join and when do we use it.
7. What are the conditions to use merge joins in Pig
8. How to implement customized loader
9. Can we store the data in partitioned manner? If yes, How?
10. How to load Hbase/Hive table in Pig
11. How to load partitioned hive table in pig
12. Which of the below efficient?

           i. A = load ' abc.txt' using PigStorage(",");
              B = Filter A by $0>10 and $1>15;
           
           ii. A = load ' abc.txt' using PigStorage(",");
               B = Filter A by $0>10;
               C = Filter B by $1>15;

Thursday, 23 April 2015

Hadoop Interview Questions 2015 - Part 2

Explain Build in sort mechanism to sort a file with single column with Billion records.

Default sort mechanism used in MapReduce?

How does a data of 200MB will be read by Mapper when block size is 64MB?

Explain secondary sort,Map&Reduce side joins

How to specify string as delimiter in hive?

Working with Counters: https://www.mapr.com/blog/managing-monitoring-and-testing-mapreduce-jobs-how-work-counters#.VU62SflVikohttps://www.mapr.com/blog/managing-monitoring-and-testing-mapreduce-jobs-how-work-counters#.VU62SflViko

Wednesday, 22 April 2015

Hadoop Interview Questions 2015 - Part 1



Few more questions.. Happy Reading



1.Explain how Hadoop is different from other parallel computing solutions.

2.What are the modes Hadoop can run in?

3.What is a NameNode and what is a DataNode?

4.What is Shuffling in MapReduce?

5.What is the functionality of Task Tracker and Job Tracker in Hadoop? How many instances of a Task Tracker and Job Tracker can be run on a single Hadoop Cluster?

6.How does NameNode tackle DataNode failures?

7.What is InputFormat in Hadoop?

8.What is the purpose of RecordReader in Hadoop?

9.Why can't we use Java primitive data types in Map Reduce?

10.Explain how do you decide between Managed & External tables in hive

11.Can we change the default location of Managed tables

12.What are the points to consider when moving from an Oracle database to Hadoop clusters? How would you decide the correct size and number of nodes in a Hadoop cluster?

13.If you want to analyze 100TB of data, what is the best architecture for that?

14.What is InputSplit in MapReduce?

15 In Hadoop, if custom partitioner is not defined then, how is data partitioned before it is sent to the reducer?

16.What is replication factor in Hadoop and what is default replication factor level Hadoop comes with?

17.What is SequenceFile in Hadoop and Explain its importance?

18.What is Speculative execution in Hadoop?

19.What are the factors that we consider while creating a hive table

20.What are the compression techniques and how do you decide which one to use

21.Co group in Pig

22.If you are the user of a MapReduce framework, then what are the configuration parameters you need to specify?

23.How do you benchmark your Hadoop Cluster with Hadoop tools?

24.Explain the difference between ORDER BY and SORT BY in Hive?

25.What is WebDAV in Hadoop?

26.How many Daemon processes run on a Hadoop System?

27.Hadoop attains parallelism by isolating the tasks across various nodes; it is possible for some of the slow nodes to rate-limit the rest of the program and slows down the program. What method Hadoop provides to combat this?

28.How are HDFS blocks replicated?

29.What will a Hadoop job do if developers try to run it with an output directory that is already present?

30.What happens if the number of reducers is 0?

31.What is meant by Map-side and Reduce-side join in Hadoop?

32.How can the NameNode be restarted?

33.How to include partitioned column in data - Hive

34.What hadoop -put command do exactly

35.What is the limit on Distributed cache size?

36.Handling skewed data

37.When doing a join in Hadoop, you notice that one reducer is running for a very long time. How will address this problem in Pig?

38.How can you debug your Hadoop code?

39.What is distributed cache and what are its benefits?

40.Why would a Hadoop developer develop a Map Reduce by disabling the reduce step?

41.Explain the major difference between an HDFS block and an InputSplit.

42.Are there any problems which can only be solved by MapReduce and cannot be solved by PIG? In which kind of scenarios MR jobs will be more useful than PIG?

43.What is the need for having a password-less SSH in a distributed environment?

44.Give an example scenario on the usage of counters.

45.Does HDFS make block boundaries between records?

46.What is streaming access?

47.What do you mean by “Heartbeat” in HDFS?

48.If there are 10 HDFS blocks to be copied from one machine to another. However, the other machine can copy only 7.5 blocks, is there a possibility for the blocks to be broken down during the time of replication?

49.What is the significance of conf.setMapper class?

50.What are combiners and when are these used in a MapReduce job?

51.What are the Different joins in hive?

52.Explain about SMB join in Hive

53.Which command is used to do a file system check in HDFS?

54.Explain about the different parameters of the mapper and reducer functions.

55.How can you set random number of mappers and reducers for a Hadoop job?

56.Did you ever built a production process in Hadoop? If yes, what was the process when your Hadoop job fails due to any reason? (Open Ended Question

57.Explain about the functioning of Master Slave architecture in Hadoop?

58.What is fault tolerance in HDFS?

59.Give some examples of companies that are using Hadoop architecture extensively.

60.How does a DataNode know the location of the NameNode in Hadoop cluster?

61.How can you check whether the NameNode is working or not?

62.Explain about the different types of “writes” in HDFS.


Hope this helps!

Saturday, 9 August 2014

Hadoop Interview Questions

Hi Reader,

Below are the questions that I faced in one of the recent interviews. Hope it helps 

Hadoop Framework:
  1. What is the difference between existing file system and HDFS? why do we need HDFS?
  2. What are the different modes?
  3. Configuration files and their properties
  4. Where do we set Name node, Data node, task tracker address/location?
  5. How many instances of Job tracker runs on a cluster?
  6. What is the difference b/w job and a task?
  7. How job tracker manages the jobs?
  8. How many task trackers exist on a data node?
  9. What happens if a job tracker fails?
  10. What happens if the Name node fails?
  11. How Secondary Name node will get the data present in Name node?
Map Reduce:

  1. Word count example flow/ Map reduce job flow?
  2. What are the phases of reducer?
  3. What is speculative execution?
  4. If two instances of same mapper gets completed at same time, what are the factors that job tracker consider in the selection of completed task?
  5. What do you mean by combiner?
  6. Where do we need to use combiner ?
  7. Which class/Interface will be used to write Combiner?
  8. What is partitioning?
  9. How to implement customized partition method?
  10. What is Map/Reduce side join? when do we go for it?Adv and disadvantages ?
  11. What is distributed cache?
Hive:
  1. When is Hive used?
  2. How to change the location of schema while creating it?
  3. What are the different properties that can be set while defining schema?
  4. What are the types of partitions?
  5. Explain a Scenario where we need a partition.
  6. What is bucketing?
  7. What are the properties that can be set in the Hive query?
  8. What does explain plan contain? 
  9. How the number of mappers and reducers will be decided in a Hive query? Example??
  10. How many number of Map reduce jobs will be created for a join query on 3 tables by same key?
  11. What are the properties to be set for query optimization?
  12. What is AVRO?
  13. Write a query to find the top 2nd student details based on his marks
  14. Write a query to filter all the duplicate records
  15. Methods to implement for an UDF?? 
  16. Commands  to run before using  UDF in a Hive query?
  17. Explain the factors to be considered in Schema/Table design
PIG:
  1. What is pig? when do we use it?
  2. What is the difference b/w Hive and Pig? when to use Pig and when to use Hive?
  3. Different joins supported by PIG?
  4. How to limit number of records?
  5. Explain foreach.
  6. what is co-group?
  7. what is bag? can a bag contain duplicate values?
  8. How to find length of a column value in Pig?
  9. Explain UDF implementation.
  10. How to define constant in pig script?
Other:
  1. What is oozie? when do we use it? 
  2. Give an example oozie program/script/code
  3. When to use Sqoop? architecture of sqoop? 
  4. Is Sqoop Map only job or Map Reduce job?

Please add more questions in comment if you have. All the very best :)

The Mindset Behind Reliable Data Systems I’ve been in data engineering long enough to see the stack change many times over. Tools come and g...