About This Course
Pass the DP-750 Azure Databricks Data Engineer Associate Exam | PySpark, Lakeflow Pipelines, CI/CD, Lakehouse, ABAC, Job
What you'll learn:
- Understand the Azure Databricks Lakehouse architecture and its components including Delta Lake, Unity Catalog, Metastore, Volumes, Managed and External Tables
- Build & Deploy Declarative Automation Bundles (DABs) for CI/CD workflows using Databricks CLI, Databricks Git Folders and Azure DevOps
- Configure different compute types and performance settings including node count, autoscaling, termination, pooling with Photon engine, and cluster policies
- Master data ingestion with Lakeflow Connect, Notebooks, Azure Data Factory, from various sources including Azure SQL, Data Lake, REST APIs, and Azure Event Hubs
- Implement data transformation on both batch and streaming data using PySpark and SparkSQL. Handle Duplicates, Nulls, Filter, Joins, Unions, Except, Pivot, Merge
- Build Lakeflow Declarative Pipelines with Streaming Tables and Materialized Views to create STAR Schema, Slowly Changing Dimensions (SCD Type) with Expectations
- Orchestrate and schedule jobs and workflows using Databricks Jobs and Azure Data Factory. Implement error handling, retries, repair, restart, stop and alerts
- Explore Databricks SQL Warehouse and learn how to work with Query Parameters, Query Caching, Query Snippets, SQL Alerts, AI/BI Genie
- Optimize with caching, partitioning, Z-ordering, Liquid Clustering, VACUUM, Deletion Vector. Tackle Spilling, Skewing, and Shuffle issues using DAG and Spark UI
- Secure and govern data with RLS, Data Masking, ABAC, Azure Key Vaults, Service Principals. Manage data lineage, access control, users & groups, retention policy
- Implement monitoring and audit logging with Databricks, Azure Monitor, and Azure Log Analytics. Enable Delta Sharing to manage data sharing with external orgs