Java Hyd Team
746 subscribers
986 photos
39 videos
670 files
690 links
https://teamhydteam.my.canva.site/

Can visit us on our website 😊
Still working on it 😊
Download Telegram
@Round 1 --

Very crucial and nervous

Current Project Explanation.

Any Major issues resolved in the current project.

What all optimization techniques have you used in your project

Hadoop Architecture

Three SQL questions majorly on joins, subqueries, Group By, inline view, with Clause, timestamp.

Coding Questions Pyspark or scala Spark.

What are different constraints in sql and why do they different?

Round 2 --

Mostly focused on Scenario based questions. Those include--

How do you increase mappers?

What if Running Job fails in sqoop?

How do you update with Latest data on Sqoop? What you do if data skewing happens?

What kind of file formats do you use? and where?

Why we need RDD?

What is the use driver in spark?

How do you tackle Memory exceptions errors in spark?

What kind of join do you use if one partition has

have less data?

Why do we need to use containers in spark?

more data and others

What happens if we increase More partitions in spark?

Write down command to increase no.of partitions?

Write down UDF query?

@Round 3 --

Important to get through or rejected

Architecture of MapReduce?

What is Outliers in MapReduce?

What is Partition, shuffle & sort?

What is block report in Hadoop?

Partition By Vs Bucketing in Hive? Difference between Cassandra and HBase?

What is Catalyst optimizer?

Explain the ETL tools have you used?

Explain basic cloud concepts and more questions on specific cloud we

worked?

DataFrame Vs Dataset? Explain Broadcast join?

Hope this helps everyone who is giving or about to give interviews.
Message @stockkida for more 😊😊