Databricks Trainer-Dutch

Phoenix Technologies

Phoenix Technologies logo
Brussel, Brussel-426

Experience

7-15 Yrs

Salary

Not disclosed

Job type

Full Time

Work mode

On-site

Skills

  • Databricks

Job description

We would like the training to cover the following topics:

• Databricks Functionality

o Databricks Unity Catalog (specifically including the Lineage View)

o Using Notebooks – %md | %sql | %python

o The MS Word-like look and feel of Notebooks

o Notebooks Revision History functionality

o Notebooks Export/Import functionality

o Visualization capabilities in Databricks (both visuals within a cell and separate dashboards)

o Using Genie

o Defining temporary views to build upon in subsequent commands

o Reading Excel files and writing them to a temporary view

o Interaction with Power BI – How to load Databricks output datasets into Power BI

o Delta Lake storage and partitioning

o Performance analysis – identifying the steps the cluster executes and pinpointing where slowdowns occur o Useful PySpark libraries

▪ Installing/importing libraries

▪ PySpark DataFrame and SQL library (pyspark.sql) and scenarios where

this is preferred over SQL

o Advanced SQL: Nested queries, JOIN clause conditions, CASE WHEN, querying

metadata (Information Schema), VERSION AS OF

o Using parameters in Databricks

o Scheduling workflows in Databricks

o Serverless vs. Dedicated Clusters -> Which to use when?

o …

Internal

• Client Data and Information Platform:

o Architecture diagram (bronze-silver-gold)

o Domain catalogs and recurring schemas

o Figures regarding current usage within Client and/or the number of tables

on Databricks

o Data Vault Modeling – ref. our BDV (Business Data Vault) schema in the silver

layer

o Dimensional Modeling – ref. our DWH (Data Warehouse) schema in the gold

layer

o Consulting the Repos logic: 'federated' code that feeds the various layers

(bronze-silver-gold) on the platform

o Where to store self-service tables and logic (analysis_ungoverned

and playground schemas for data, and folders for logic)

We expect a significant number of exercises and case studies relevant to Client to be

developed during the training, based on Client's internal data.

Learning objectives: Upon completion of this training, participants will be able to:

• Independently build self-service ETL processes on the Databricks platform using SQL and/or

PySpark

• Make optimal use of Databricks features.

• To design a logical data model that can be technically implemented by a more specialized Data Engineer

• To be able to develop a prototype star schema that can be technically implemented by a more specialized Data Engineer

• To schedule workflows and optimize queries

 

Skills required:

In the case of a classroom training, the teacher can be
present at the requested location at 7.30 am at the latest

The teacher can make himself available within 6 weeks to give
a training (from request for new session(s))

The trainer has demonstrable experience as an instructor

The trainer has a thorough knowledge of and experience with
the Databricks Platform

The trainer has a strong command of programming languages SQL
and Python for data analysis and machine learning on Databricks

The trainer can move to all specified locations

The trainer shows in the plan of action that he has provided
similar training courses in the past. We expect at least 2 reference examples
(where, when, content/short description, sector/customer,..

The trainer has a certification from Databricks

The trainer has experience in implementing advanced data
warehousing and data governance processes

The trainer has experience in developing and maintaining
machine learning models and data visualization tools.

The trainer has at least 2 years of experience in the energy
sector NOT Mandatory Its plus

About the company

Phoenix Technologies

phoenix-tech.eu