You’ve probably seen this situation if you’ve worked on a growing data platform: You’re on a dashboard when it shows the wrong number and you don’t know where things went wrong.
You unlock the dashboard, then the Gold table, then the transformation job, and next the Silver table. In no time at all, you’re bouncing between notebooks, SQL queries, pipelines and source systems to get to the point of one broken data dependency.
Data lineage in Databricks can solve that problem.
To visualize the movement of your data from its source to downstream tables, columns, jobs, dashboards, or other assets, instead of manually tracing your data environment. Databricks Unity Catalog takes lineage out of your data governance process and makes it an integral part of it.
So, what is Databricks data lineage, and how can you use it effectively? Let’s divide it up.
What is Data Lineage in Databricks?
Data lineage in Databricks is a tool that provides information on the origin of data, how it’s being transformed, and where it’s used after that transformation.
Consider a standard Databricks set-up:
Source System → Bronze → Silver → Gold → Dashboard
There could be customer information in your source system. That data is fed into a Bronze table, cleaned and transformed in Silver, aggregated in Gold and then pushed to a BI dashboard. Now, suppose that someone modifies the customer_id field in the Silver layer.
What depends on it?
Lineage is the place it comes in handy. Instead of manually searching the environment, you can follow the upstream and downstream dependencies. Databricks preserves supported lineage in Unity Catalog, including lineage down to the column level.
How Does Databricks Data Lineage Work?
It’s tightly coupled with Databricks’ governance layer, Unity Catalog. Supported queries run in Databricks preserve the relationships between the Data Assets used in the query. Those relationships can then be investigated using Catalog Explorer.
Suppose that your Gold table has:
- customer_id
- order_count
- total_revenue
The total_revenue attribute could be a result of computations on one or more of the preceding attributes. In addition to table-level lineage, you can explore those relationships, rather than just at the table level.
This separation is important when doing real projects. It helps to have knowledge that sales_gold requires sales_silver. A good understanding of which columns are providing data for total_revenue is more valuable when debugging a report that’s critical to the business.
Databricks also provides external lineage, in which you can model relationships to systems outside Databricks. Sources from Salesforce, MySQL, or another application can be linked to downstream tools like Tableau or Power BI to get the big picture of your data flow.
What can Databricks Data Lineage Track?
There are multiple ways in which your lineage graph can be used to gain visibility into your data environment.
Tables and Views
One can visualize the flow of information to datasets from downstream tables and views. This provides more visibility into the data pipeline dependency.
Columns
Column-level lineage further analyzes the data. You can delve into how each field is calculated, based on the data in the upstream fields.
Notebooks and Queries
Notebooks and queries can add context if there is a need for it, when trying to understand what was happening in a transformation.
Jobs and Pipelines
Jobs & pipelines are identified that are connected to your dataset. This can be especially helpful when looking into transformations that didn’t happen or transformations that have failed.
Dashboards
Lineage is useful for identifying those relationships if the data set that a business dashboard relies on is set to be changed. In addition to supporting lineage for tables and pipelines, Databricks also supports lineage for some AI and ML assets governed by Unity Catalog.
Why Is Data Lineage Important in Databricks?
As for me, in my experience, the more important it is that lineage comes in when things change. If you have a Gold table, and you’d like to remove an old column, what should you do? If there is no lineage, you may spend hours combing through hundreds of SQLs and notebooks to see if you’ve overlooked a dependency somewhere. Lineage allows for researching downstream relationships first.
1. You Can Perform Impact Analysis
You can detect potential downstream assets that might rely on a table, view, or column before modifying the asset. That provides your team with a much safer alternative to planning schema changes.
2. You Can Troubleshoot Faster
Suppose there’s a financial dashboard that suddenly shows the wrong amount of revenue. You can choose to investigate back up the lineage graph from the output, rather than from the dashboard and making a guess.
3. You Strengthen Data Governance
Good Databricks data governance isn’t just about permissions. You also want to know how information flows in your platform. This is the kind of visibility Lineage provides you.
4. You Build More Trust in Your Data
Your analysts and data consumers have more context to use before they use it when you can trace a dataset back to its source and understand the transformations that happened.
5. You Can Understand External Dependencies
When you have an architecture that consists of systems beyond Databricks, external lineage can help you document that relationship rather than viewing Databricks as an isolated environment.
How to View Data Lineage in Databricks
For users of Catalog, lineage is available in Catalog Explorer. The basic workflow is the following:
- Open a Catalog in your Databricks workspace.
- Find the tables and/or views you wish to explore.
- Click on the Lineage tab.
- Select See Lineage Graph.
- Extend graph to explore upstream and downstream relationships.
- Choose a column to explore the relationships between columns.
- You can also filter related assets like notebooks, jobs, pipelines and queries.
An important note: What you can see is related to what permissions you have. Access controls are also applied to lineage by Unity Catalog, meaning that restricted objects might not be entirely visible to you.
What Are the Limitations of Databricks Data Lineage?
This is where it’s time to be practical.
Databricks data lineage isn’t a magical record of all the data movements your organization carries out. Lineage coverage is dependent upon the configuration of your workloads and data. For instance, Databricks notes the lack of support for RDDs, renamed dBs, global temporary views, and certain column-level lineage scenarios, among other things.
It is also possible to restrict column-level lineage if the source or target data is accessed by a file path instead of a registered table name. The introduction of additional constraints may be caused by UDFs, since Databricks may store table-level relationships but not the column-to-column relationships.
So, don’t treat the lineage graph as your only source of documentation. Use it in conjunction with good data architecture, metadata, governance policies and data quality practices.
Best Practices for Using Data Lineage in Databricks
To get lineage to be useful on a daily basis in your data engineering process, make it a part of your workflow. You should:
- Use Unity Catalog for centralized governance.
- Maintain uniform naming for catalog, schema, tables and columns.
- Properly register important datasets.
- Perform schema changes after reviewing lineage.
- Use lineage when investigating data quality problems.
- Document important external dependencies.
- Set permissions on all assets of data.
- Integrate lineage with data governance/metadata management.
The most common error is to consider lineage as a backup only when a pipeline fails. When lineage is integrated with the way you design, build, and govern your Databricks environment, your team gets much more value.
Final Thoughts
Data lineage in Databricks provides visibility into data flows throughout your platform. With investigating relationships in Unity Catalog and Catalog Explorer you can now find out instead of trying to guess what data, column, job or dashboard is related to another.
That visibility is even more crucial as your Databricks environment expands. It can support you to troubleshoot quicker, appraise the effect of modifications, enhance data governance, and offer your teams’ trust in the data they are using.
For more information on the Databricks architecture, data engineering, data governance, and solutions, you can further explore the top Databricks consulting companies and connect with Databricks partenrs on bricksperformers