What this is
A Databricks source lets Zeotap CDP ingest data from a table in your Databricks workspace on a defined refresh schedule. Two connection mechanisms are supported — JDBC, which reads through a Databricks SQL warehouse, and Job Based, which runs a Databricks job to move the data. The Job-based approach is the recommended path for tables above 1 million records, since JDBC may encounter issues at that data volume. Data ingested from Databricks lands in either the Customer Data or Non Customer Data entity in Zeotap CDP.Prerequisites
Required permissions in Databricks
The user or service principal used for authentication must have:- Unrestricted Cluster Access, or a cluster policy that permits creating the required node-type clusters.
- SELECT permission on the required tables.
- USE SCHEMA permission on the relevant schema.
Components needed for every Databricks source
Obtain these four values from your Databricks account before you create the source in Zeotap CDP.Databricks Host
The Databricks Host is the unique URL assigned to your Databricks workspace. You can find it in the URL for the Databricks account.
Catalog Name
In Databricks, a catalog acts like a folder system. It organises schemas and databases into a hierarchy, allowing users to group tables and views logically.
Schema Name
In Databricks, a schema is essentially a database that contains tables. It serves as a container for organising and managing related tables, providing a structured way to store and retrieve data.
Table Name
This is the name of the table within the schema that you want to ingest.
Additional components for the JDBC mechanism
- HTTP Path of a Databricks SQL warehouse.
- Client ID and Client Secret from an OAuth service principal.
- Partition Column or Unique Column (see Partition Column and Unique Column).
1. Create a Service Principal
2. Assign Workspace-Level Permissions to the Service Principal
3. Create an OAuth Secret for the Service Principal
Databricks Client ID and Client Secret
Before you can use OAuth to authenticate to Databricks, you must first create an OAuth Client Secret, which is used to generate OAuth access tokens. A service principal can hold up to five OAuth secrets. Account admins and workspace admins can create an OAuth secret for a service principal. For the specific procedure to generate the Client ID and Client Secret, see step 3 of the Databricks OAuth documentation. To create a service principal and generate its Client ID and Client Secret, go to Identity and Access → Service Principals → Add Service Principal → Generate Secret in your Databricks workspace:Open Service Principals

Add the service principal

Generate the secret

HTTP Path
To locate the HTTP Path in Databricks, perform the following steps:Open SQL Warehouses

Select the SQL warehouse

Open Connection details

Copy the HTTP Path
Partition Column and Unique Column
A Partition Column is a key used to organise large tables into smaller, logical groups. For example, data can be partitioned by date or region, so queries targeting a specific time range or location run much faster. If no Partition Column exists on the source table, use any Unique Column instead — a column whose values are different in every row. When only a Unique Column is available, Zeotap CDP:- Creates a temporary view and runs the read queries over that view.
- Drops the view once all the data has been fetched.
Additional components for the Job-based mechanism
For the Job Based approach, select an Auth Type and provide the corresponding credentials. The supported auth types are PAT Token and Service Principal.PAT Token
A Personal Access Token (PAT) can be generated in Databricks under Settings → User setting → Developer → Generate new token.
- jobs
- secrets
- sql
- unity-catalog
- workspace
- Creating and running jobs
- Managing secret scopes
- Accessing tables via SQL and Unity Catalog
- Creating notebooks and workspace artefacts
Service Principal
Provide the Client ID and Client Secret of your Databricks service principal. These are the same credentials used for the JDBC approach — see Databricks Client ID and Client Secret for how to create the service principal and generate the OAuth secret.Choose a connection mechanism
Configure the source in Zeotap CDP
Once you have gathered the Databricks components, create the source in the Zeotap CDP App.Open the Sources application
Select Data Warehouse as the category
.png?fit=max&auto=format&n=ROPrHg77hrORMuiL&q=85&s=146e94f6110f76970d52baf2ab1e964b)
Select Databricks as the data source

Name the source and pick the region
Set the Refresh Frequency
- Sync once
- Every hour
- Every 3 hours
- Every 6 hours
- Every 12 hours
- Daily
- Weekly
- Monthly
Enter host, catalog, schema, and table

Choose the data entity
Configure Delta Queries Selection

- true — only new and updated values, based on the timestamp in the delta column, are fetched. Delta Column Name and Delta Column Data Type become active: enter the name of the column to fetch data from, and choose the time increment for fetching the data from that column.
- false — no additional selections are needed. Zeotap CDP fetches all existing data from the table on every run.

Select the connection mechanism

Enter the credentials for the chosen mechanism
- Client ID
- Client Secret
- HTTP Path
- Partition Table Details (Partition Column Name, Partition Column Type)
- Unique Column Name




Select the fields to ingest
Create the source
Field-level requirements
The fields a Databricks source configuration collects:Verify the source is integrated
Open the source from the Sources listing and go to the IMPLEMENTATION DETAILS tab. A successfully created source shows all the relevant information about the created source on this tab and appears on the Sources listing with the Integrated status.
Troubleshooting
If the source does not reach Integrated, or the first sync does not produce data, check the conditions below.- The source name and connection mechanism (JDBC or Job Based), and for Job Based the Auth Type.
- The refresh frequency.
- The timestamp of the last completed sync (if any).
- The Databricks catalog, schema, and table names.
FAQ
Which connection mechanism should I pick — JDBC or Job Based?
Which connection mechanism should I pick — JDBC or Job Based?
Which Auth Type should I use for the Job Based mechanism?
Which Auth Type should I use for the Job Based mechanism?
What if the source table has no partition column?
What if the source table has no partition column?
How long does the first sync take?
How long does the first sync take?