all companies

Databricks data interview guide

Lakehouse and pipeline work — Spark-style transforms and data modeling.

typical interview loop

  1. 1.Recruiter screen
  2. 2.Technical screen: SQL plus a Python data-manipulation exercise
  3. 3.Onsite: data modeling / pipeline design, coding, behavioral
  • › pandas / PySpark reshaping and joins are central.
  • › Be ready to discuss partitioning and incremental loads.

most-tested topics

ratios
3
groupby
2
JOIN
2
GROUP BY
2
CASE
2
graphs
2

Based on 18 reconstructed Databricks-style questions.

practice path

0 free · 18 Pro
  1. 1Average runtime per jobpandasEasy
  2. 2Cluster compute hourssqlMedium
  3. 3Compact small filespythonHard
  4. 4Cost per successful runsqlHard
  5. 5Daily pipeline healthpandasHard
  6. 6Delta size per catalogsqlEasy
  7. 7Job failure ratesqlMedium
  8. 8Parse a cluster specpythonEasy
  9. 9Recursive lineage descendantssqlHard
  10. 10Retry hotspotspandasMedium
  11. 11Retry with exponential backoffpythonMedium
  12. 12Run duration trend per jobsqlHard
  13. 13Runs per jobsqlEasy
  14. 14Small file ratio per catalogpandasMedium
  15. 15Spot versus on-demand spendsqlMedium
  16. 16Stale vacuum auditsqlMedium
  17. 17Tables needing compactionsqlMedium
  18. 18Topological order of a pipelinepythonHard

Unlock all 18 Databricks questions with Pro.

upgrade