Databricks Trainer-Dutch
Phoenix Technologies
Experience
7-15 Yrs
Salary
Not disclosed
Job type
Full Time
Work mode
On-site
Skills
- Databricks
Job description
We would like the training to cover the following topics:
• Databricks Functionality
o Databricks Unity Catalog (specifically including the Lineage View)
o Using Notebooks – %md | %sql | %python
o The MS Word-like look and feel of Notebooks
o Notebooks Revision History functionality
o Notebooks Export/Import functionality
o Visualization capabilities in Databricks (both visuals within a cell and separate dashboards)
o Using Genie
o Defining temporary views to build upon in subsequent commands
o Reading Excel files and writing them to a temporary view
o Interaction with Power BI – How to load Databricks output datasets into Power BI
o Delta Lake storage and partitioning
o Performance analysis – identifying the steps the cluster executes and pinpointing where slowdowns occur o Useful PySpark libraries
▪ Installing/importing libraries
▪ PySpark DataFrame and SQL library (pyspark.sql) and scenarios where
this is preferred over SQL
o Advanced SQL: Nested queries, JOIN clause conditions, CASE WHEN, querying
metadata (Information Schema), VERSION AS OF
o Using parameters in Databricks
o Scheduling workflows in Databricks
o Serverless vs. Dedicated Clusters -> Which to use when?
o …
Internal
• Client Data and Information Platform:
o Architecture diagram (bronze-silver-gold)
o Domain catalogs and recurring schemas
o Figures regarding current usage within Client and/or the number of tables
on Databricks
o Data Vault Modeling – ref. our BDV (Business Data Vault) schema in the silver
layer
o Dimensional Modeling – ref. our DWH (Data Warehouse) schema in the gold
layer
o Consulting the Repos logic: 'federated' code that feeds the various layers
(bronze-silver-gold) on the platform
o Where to store self-service tables and logic (analysis_ungoverned
and playground schemas for data, and folders for logic)
We expect a significant number of exercises and case studies relevant to Client to be
developed during the training, based on Client's internal data.
Learning objectives: Upon completion of this training, participants will be able to:
• Independently build self-service ETL processes on the Databricks platform using SQL and/or
PySpark
• Make optimal use of Databricks features.
• To design a logical data model that can be technically implemented by a more specialized Data Engineer
• To be able to develop a prototype star schema that can be technically implemented by a more specialized Data Engineer
• To schedule workflows and optimize queries
Skills required:
In the case of a classroom training, the teacher can be
present at the requested location at 7.30 am at the latest
The teacher can make himself available within 6 weeks to give
a training (from request for new session(s))
The trainer has demonstrable experience as an instructor
The trainer has a thorough knowledge of and experience with
the Databricks Platform
The trainer has a strong command of programming languages SQL
and Python for data analysis and machine learning on Databricks
The trainer can move to all specified locations
The trainer shows in the plan of action that he has provided
similar training courses in the past. We expect at least 2 reference examples
(where, when, content/short description, sector/customer,..
The trainer has a certification from Databricks
The trainer has experience in implementing advanced data
warehousing and data governance processes
The trainer has experience in developing and maintaining
machine learning models and data visualization tools.
The trainer has at least 2 years of experience in the energy
sector NOT Mandatory Its plus
About the company
Phoenix Technologies
phoenix-tech.eu