Skip to main content
This page is the complete reference for the Databricks source — components, prerequisites, and full setup. For a shorter setup-only walkthrough that skips the reference material, see Set Up a Databricks Source.

What this is

A Databricks source lets Zeotap CDP ingest data from a table in your Databricks workspace on a defined refresh schedule. Two connection mechanisms are supported — JDBC, which reads through a Databricks SQL warehouse, and Job Based, which runs a Databricks job to move the data. The Job-based approach is the recommended path for tables above 1 million records, since JDBC may encounter issues at that data volume. Data ingested from Databricks lands in either the Customer Data or Non Customer Data entity in Zeotap CDP.

Prerequisites

Setting up a Databricks source requires access to your Databricks workspace with permission to create a service principal, generate the credentials for your chosen mechanism (OAuth Client ID and Client Secret, or a personal access token), and read the target table. You also need access to the Sources application in the Zeotap CDP App.

Required permissions in Databricks

The user or service principal used for authentication must have:
  • Unrestricted Cluster Access, or a cluster policy that permits creating the required node-type clusters.
  • SELECT permission on the required tables.
  • USE SCHEMA permission on the relevant schema.

Components needed for every Databricks source

Obtain these four values from your Databricks account before you create the source in Zeotap CDP.

Databricks Host

The Databricks Host is the unique URL assigned to your Databricks workspace. You can find it in the URL for the Databricks account.
The Databricks workspace URL in the browser address bar, which is the Databricks Host value

Catalog Name

In Databricks, a catalog acts like a folder system. It organises schemas and databases into a hierarchy, allowing users to group tables and views logically.
Catalog listing in the Databricks workspace

Schema Name

In Databricks, a schema is essentially a database that contains tables. It serves as a container for organising and managing related tables, providing a structured way to store and retrieve data.
Schema listed under a catalog in the Databricks workspace

Table Name

This is the name of the table within the schema that you want to ingest.
Table listed under a schema in the Databricks workspace

Additional components for the JDBC mechanism

  • HTTP Path of a Databricks SQL warehouse.
  • Client ID and Client Secret from an OAuth service principal.
  • Partition Column or Unique Column (see Partition Column and Unique Column).
Before entering JDBC credentials in Zeotap CDP, complete the OAuth setup in Databricks by following the Databricks OAuth M2M documentation:

1. Create a Service Principal

2. Assign Workspace-Level Permissions to the Service Principal

3. Create an OAuth Secret for the Service Principal

Databricks Client ID and Client Secret

Before you can use OAuth to authenticate to Databricks, you must first create an OAuth Client Secret, which is used to generate OAuth access tokens. A service principal can hold up to five OAuth secrets. Account admins and workspace admins can create an OAuth secret for a service principal. For the specific procedure to generate the Client ID and Client Secret, see step 3 of the Databricks OAuth documentation. To create a service principal and generate its Client ID and Client Secret, go to Identity and Access → Service Principals → Add Service Principal → Generate Secret in your Databricks workspace:
1

Open Service Principals

Go to Identity and Access → Service Principals and click Add service principal.
Identity and Access settings in Databricks with the Add service principal action
2

Add the service principal

Select an existing service principal from your account, or create a new one, then click Add service principal.
Selecting an existing service principal or creating a new one before clicking Add service principal
3

Generate the secret

Open the service principal, go to the Secrets tab, and click Generate secret. Copy the generated Client ID and Client Secret — the secret value is shown only once.
Secrets tab of a Databricks service principal with the Generate secret action

HTTP Path

To locate the HTTP Path in Databricks, perform the following steps:
1

Open SQL Warehouses

Log in to your Databricks instance and, under SQL, click SQL Warehouses.
SQL Warehouses entry under SQL in the Databricks navigation
2

Select the SQL warehouse

Click the desired SQL warehouse. If no SQL warehouse exists, create one by clicking Create SQL Warehouse.
SQL warehouse list in Databricks with the Create SQL Warehouse option
3

Open Connection details

On the SQL warehouse summary page, go to the Connection details tab.
Connection details tab on the Databricks SQL warehouse summary page
4

Copy the HTTP Path

Copy the HTTP Path value from the displayed information.

Partition Column and Unique Column

A Partition Column is a key used to organise large tables into smaller, logical groups. For example, data can be partitioned by date or region, so queries targeting a specific time range or location run much faster. If no Partition Column exists on the source table, use any Unique Column instead — a column whose values are different in every row. When only a Unique Column is available, Zeotap CDP:
  1. Creates a temporary view and runs the read queries over that view.
  2. Drops the view once all the data has been fetched.

Additional components for the Job-based mechanism

For the Job Based approach, select an Auth Type and provide the corresponding credentials. The supported auth types are PAT Token and Service Principal.

PAT Token

A Personal Access Token (PAT) can be generated in Databricks under Settings → User setting → Developer → Generate new token.
Generate new token under Developer in the Databricks user settings
Required PAT scopes. When using PAT authentication, the token must carry the following scopes:
  • jobs
  • secrets
  • sql
  • unity-catalog
  • workspace
These scopes are required for:
  • Creating and running jobs
  • Managing secret scopes
  • Accessing tables via SQL and Unity Catalog
  • Creating notebooks and workspace artefacts

Service Principal

Provide the Client ID and Client Secret of your Databricks service principal. These are the same credentials used for the JDBC approach — see Databricks Client ID and Client Secret for how to create the service principal and generate the OAuth secret.

Choose a connection mechanism

Use Job Based for tables above 1 million records. Use JDBC for smaller tables and ad-hoc use cases where a SQL warehouse is already provisioned.

Configure the source in Zeotap CDP

Once you have gathered the Databricks components, create the source in the Zeotap CDP App.
1

Open the Sources application

In the Zeotap CDP App, open the Sources application under Integrate, then click CREATE SOURCE.
2

Select Data Warehouse as the category

Under Category, choose Data Warehouse.
Choosing Data Warehouse as the source category
3

Select Databricks as the data source

Click Databricks as the Data Source.
Selecting Databricks as the data source
4

Name the source and pick the region

Enter a short, descriptive Source Name and choose the Region of upload.
5

Set the Refresh Frequency

Choose the Refresh Frequency from the drop-down. The first sync runs when the source is created; subsequent syncs follow the frequency you select. Supported frequencies:
  • Sync once
  • Every hour
  • Every 3 hours
  • Every 6 hours
  • Every 12 hours
  • Daily
  • Weekly
  • Monthly
For Daily, Weekly, and Monthly, set Sync Time (the exact time for the sync to occur), Sync Period (whether the Sync Time is AM or PM), and — for Monthly — the Monthly Sync date (the day of the month on which the sync should run).
6

Enter host, catalog, schema, and table

Enter your Databricks workspace URL in the Databricks Host field, followed by Catalog Name, Schema Name, and Table Name.
Databricks Host, Catalog Name, Schema Name, and Table Name fields
7

Choose the data entity

Under Data Entity, select Customer Data or Non Customer Data depending on the type of data you are ingesting. For the distinction, see Supported Data Entities.
8

Configure Delta Queries Selection

Under Delta Queries Selection, decide whether to consider deltas (data additions and changes) in the table based on a timestamp column. Choose true or false:
Delta Queries Selection with true and false options
  • true — only new and updated values, based on the timestamp in the delta column, are fetched. Delta Column Name and Delta Column Data Type become active: enter the name of the column to fetch data from, and choose the time increment for fetching the data from that column.
  • false — no additional selections are needed. Zeotap CDP fetches all existing data from the table on every run.
Delta Column Name and Delta Column Data Type fields shown when Delta Queries Selection is true
9

Select the connection mechanism

Select the mechanism used to pull the data. The currently supported options are JDBC and Job Based.
Connection mechanism selection showing JDBC and Job Based options
10

Enter the credentials for the chosen mechanism

If you selected JDBC, enter the following values (see Additional components for the JDBC mechanism for how to obtain them):
  • Client ID
  • Client Secret
  • HTTP Path
  • Partition Table Details (Partition Column Name, Partition Column Type)
  • Unique Column Name
JDBC credential fields for the Databricks source
Remaining JDBC credential fields on the Databricks source configuration screen
If you selected Job Based, choose the Auth Type and enter the corresponding credentials (see Additional components for the Job-based mechanism).If the Auth Type is PAT Token, enter the Access Token:
Job Based mechanism with Auth Type set to PAT Token and the Access Token field
If the Auth Type is Service Principal, enter the Client ID and Client Secret:
Job Based mechanism with Auth Type set to Service Principal and the Client ID and Client Secret fields
11

Select the fields to ingest

Click Next to proceed to field selection. A list of fields is displayed: tick the columns you want to ingest, use Select All to pick every column available in your Databricks account, or search by name if you already know the field names.
12

Create the source

Click CREATE SOURCE. The source appears on the Sources listing page.

Field-level requirements

The fields a Databricks source configuration collects: Connection fields, by mechanism:

Verify the source is integrated

Open the source from the Sources listing and go to the IMPLEMENTATION DETAILS tab. A successfully created source shows all the relevant information about the created source on this tab and appears on the Sources listing with the Integrated status.
IMPLEMENTATION DETAILS tab of a created Databricks source
The first data transfer starts when the source is created; its duration depends on the volume of data in the source table.

Troubleshooting

If the source does not reach Integrated, or the first sync does not produce data, check the conditions below. If the source configuration passes every check but no data arrives after the first refresh, contact the Zeotap support team at support@zeotap.com. Include:
  • The source name and connection mechanism (JDBC or Job Based), and for Job Based the Auth Type.
  • The refresh frequency.
  • The timestamp of the last completed sync (if any).
  • The Databricks catalog, schema, and table names.

FAQ

Pick Job Based for any table above 1 million records; JDBC may encounter issues at that data volume. For smaller tables, either mechanism works — JDBC reuses an existing SQL warehouse and needs the Client ID, Client Secret, HTTP Path, and either a Partition Column or a Unique Column.
Both are supported. PAT Token needs a single Databricks personal access token carrying the jobs, secrets, sql, unity-catalog, and workspace scopes. Service Principal reuses the same Client ID and Client Secret you would generate for the JDBC mechanism, so it is the simpler choice if that service principal already exists.
Provide a Unique Column instead — a column whose values are different in every row. Zeotap CDP creates a temporary view over the table to run the read queries and drops the view after all the data is fetched.
The initial data transfer from Databricks to Zeotap CDP starts when the source is created and may take time depending on the data volume. Subsequent runs follow the Refresh Frequency and, when Delta Queries Selection is set to true, pull only the delta since the last successful sync.

Next steps

Last modified on September 24, 2026