Skip to main content

Posts

What is InfluxDB

InfluxDB  is an efficient, reliable, and schema-less time-series database that can store time-series data. It is a NoSQL database that provides high performance in terms of throughput, compression, and retention. InfluxDB can handle millions of time-stamped data points per second. InfluxDB includes support for real-time storage and analytics, IoT sensor data, and DevOps Monitoring.  Some of the essential components of InfluxDB are : Timestamp  As InfluxDB is a time series database, time is an important essence in it. It stores time in the form of timestamps in the RFC3339 UTC format, which is yyyy-mm-ddThh:mm:ssZ.  Fields  InfluxDB has a concept of Fields that has components such as Fields keys of string types which are similar to the columns in RDBMS, Fields values that are the actual measured values of any types string, float, integer, or boolean and Fields set is a combination of Fields keys and values.  Tags InfluxDB has one optional component called T...

Advanced SQL interview questions for Data Engineers

SQL is one of the most favorite topics in Data Engineering interviews because Data Engineers should not only be proficient in Programming but should be able to write simple or advanced sql queries, Data Modelling, and Pipeline Design. In this post, we will discuss the possible advanced sql questions asked in Data Engineering Interviews. Advanced SQL Interview Questions 1.  What is the difference between GROUP BY and PARTITION BY? GROUP BY PARTITION BY GROUP BY returns only one row after aggregating columns for each group. PARTITION BY gives aggregated columns for each record in the table. The number of rows in the table is reduced. The number of rows in the table remains the same. It is an aggregation function. It is an analytic function. GROUP BY does not allow to add the columns that are not a part of the GROUP BY clause. With the PARTITION BY clause, we can add any columns. 2.  How to...

What is Shuffling in Spark

Shuffling in Spark is a mechanism that Re-Distributes the data across different executors or workers in the clusters.  Why do we need to Re-Distribute the data?    A) Re-Distribution is needed when there is a need of increasing or decreasing the data partitions in the situations below: When the partitions are not sufficient enough to process the data load in the cluster When the partitions are too high in numbers that it creates task scheduling overhead and it becomes the bottleneck in the processing time. Re-Distribution can also be achieved by executing the shuffling on existing distributed data collection like RDD, DataFrames, etc by using the "Repartition" and "Coalesce" APIs in Spark. B) During Aggregation and Joins on data collection in Spark, all the data records belonging to aggregation or join should reside in the single partition and when the existing partitioning scheme doesn't satisfy this condition there is a need to re-distributing the data in in...

What is NoSQL Database?

The name NoSQL itself tells us that it is a  "non-SQL"  or  "non-relational"  database. Around 30 years back when the data used to be non-changing and smaller in size, traditional relational databases were more prominent like ORACLE, Postgres and so on which had fixed schemas. But during the last decade, the data has grown exponentially and it is also changing quickly. The traditional databases have failed to handle this BIG DATA effectively. So there was a need to introduce a database that can adapt itself to ever-changing data and that can handle the enormous size of data. And thus NoSQL databases came into the picture. Nowadays NoSQL databases have been referred to as  "Not Only SQL"  databases which mean that these databases may support SQL-like query languages and can be a part of polyglot persistent architecture along with other relational databases. The data structures used in the NoSQL database are more efficient than the data structures used by th...

Difference between Union and Union All in SQL

You might be using Union or Union All in your SQL code while doing Data Analysis or building Data Pipelines. Ever wondered what is the difference between them and how using one over another can be more efficient? Yes, there is a small yet significant difference between Union and Union All. Let's look at that by understanding each of them individually. 1. Union All  Union All basically allows you to concatenate the table that has a similar structure of tables. The important condition to have Union All of the tables is that both the tables should have the same number of columns. So when you take Union All of two tables what it does in the background is it directly joins the tables without removing duplicates or redundant records.   2. Union  Union is also similar to Union All except one difference that it removes the duplicates records before taking the Union of the tables.  There is one disadvantage of Union over Union All, that since it removes duplicated records bef...

What is CAP Theorem?

CAP Theorem states that a Distributed Database System can only have 2 out of 3 properties from Availability, Consistency, and Partition Tolerance . This means that every Big Data Engineer needs to do a trade-off between these three based on the use-case and Business requirements. It is very important for any Data engineer to understand the CAP Theorem and apply it when deciding the appropriate tools for the task in the hand. Let's discuss each of the properties in detail. 1. Availability   This condition states that every request (read/write) will get a response on Success or Failure. That means every node in the system must return a response in a reasonable amount of time. This could be only possible if the system remains operational all the time. Hence, the databases are time-independent as they should be available all the time. Therefore if any two records are added to the database we don't know which one was added first and the output could be either one of them. Now le...

Top 25 Data Engineer Interview Questions

In my last post  How to prepare for Data Engineer Interviews ,  I wrote about how one can prepare for the Data Engineer Interviews, and in this blog post, I am going to provide the  Top 25 Basic   data engineer interview questions  asked frequently and their brief answers. This is typically the first round of the Interview where the interviewer just wants to access whether you are aware of basic concepts or not and therefore you don't need to explain it in detail. Just a single statement would be sufficient. Let's get started Checkout the 5 Key Skills Data Engineers need in 2023 A. Programming  1. What is the Static method in Python? Static methods are the methods that are bound to the  Class  rather than the Class's Object. Thus, it can be called without creating objects of the class. We can just call it using the reference of the class. Also, all the objects of the class share only one copy of the static method. 2. What is a Decorator in Python?...