NYU Classifieds>NYU Online Courses>Data Management with Databricks: Big Data with Delta Lakes

Data Management with Databricks: Big Data with Delta Lakes

About this Course

In this 2-hour guided project, \"Data Management with Databricks: Big Data with Delta Lakes\" you will collaborate with the instructor to achieve the following objectives: 1-Create Delta Tables in Databricks and write data to them. Gain hands-on experience in setting up and managing Delta Tables, a powerful data storage format optimized for performance and reliability. 2-Transform a Delta table using Python and leverage SQL to query the data for creating a comprehensive dashboard. Learn how to apply Python-based transformations to Delta Tables, and use SQL queries to extract the necessary insights for building a Supply Chain dashboard. 3-Utilize Delta Lake\'s merge operation and version control capabilities to efficiently update Delta Tables. Explore the capabilities of Delta Lake\'s merge operation to perform upserts and other data updates efficiently. Additionally, learn how to leverage Delta Lake\'s built-in version control to track and access previous versions of Delta Tables as needed. Throughout a real-world business scenario, you will use Databricks to build an end-to-end data pipeline that integrates various JSON data files and applies transformations, ultimately providing valuable insights and analysis-ready data. This intermediate-level guided project is designed for data engineers who build data pipelines for their companies using Databricks. In order to be successful in this guided project, you need prior knowledge of writing Python scripts including importing libraries, setting-up variables, manipulating data frames, and using functions. You will also need to be familiar with writing SQL queries such as aggregating, filtering, and joining tables.

Created by: Coursera Project Network


Related Online Courses

This Specialization is designed for people who are new to software engineering. It\'s also for those who have already developed software, but wish to gain a deeper understanding of the underlying... more
The demand for professionals with a knowledge of artificial intelligence (AI) is on the rise. There is a revolution in the way organizations make decisions on the basis of generative AI data... more
Former U.S. Secretary of the Treasury Timothy F. Geithner and Professor Andrew Metrick survey the causes, events, policy responses, and aftermath of the recent global financial crisis.Created by:... more
The aim of this course is to introduce learners to open-source R packages that can be used to perform clinical data reporting tasks. The main emphasis of the course will be the clinical data flow... more
The Sector Investing course provides an in-depth exploration of how to allocate investments across various sectors of the economy. Investors will learn how to identify, analyze, and invest in... more

CONTINUE SEARCH

FOLLOW COLLEGE PARENT CENTRAL