Databricks data interview guide
Lakehouse and pipeline work — Spark-style transforms and data modeling.
typical interview loop
- 1.Recruiter screen
- 2.Technical screen: SQL plus a Python data-manipulation exercise
- 3.Onsite: data modeling / pipeline design, coding, behavioral
- › pandas / PySpark reshaping and joins are central.
- › Be ready to discuss partitioning and incremental loads.
most-tested topics
ratios
3
groupby
2
JOIN
2
GROUP BY
2
CASE
2
graphs
2
Based on 18 reconstructed Databricks-style questions.
practice path
0 free · 18 Pro- 1Average runtime per jobpandasEasy
- 2Cluster compute hourssqlMedium
- 3Compact small filespythonHard
- 4Cost per successful runsqlHard
- 5Daily pipeline healthpandasHard
- 6Delta size per catalogsqlEasy
- 7Job failure ratesqlMedium
- 8Parse a cluster specpythonEasy
- 9Recursive lineage descendantssqlHard
- 10Retry hotspotspandasMedium
- 11Retry with exponential backoffpythonMedium
- 12Run duration trend per jobsqlHard
- 13Runs per jobsqlEasy
- 14Small file ratio per catalogpandasMedium
- 15Spot versus on-demand spendsqlMedium
- 16Stale vacuum auditsqlMedium
- 17Tables needing compactionsqlMedium
- 18Topological order of a pipelinepythonHard
Unlock all 18 Databricks questions with Pro.
upgrade