# Data integration

> Our data integration services get the data out of the systems your business runs on, such as Business Central, Dynamics 365, Salesforce and HubSpot, and into one place where it can be joined, reported on and trusted. That place is usually Databricks. We use a managed connector where one exists and build the pipeline where none does, as with Business Central.

For: Data, finance and operations leaders whose numbers live in five systems
Canonical: https://www.techfabric.com/databricks/data-integration

---

Data integration services for Business Central, Salesforce and the rest of your systems

Business Central is where we start most often, and it's the harder case: Databricks has no managed connector for it, so the pipeline has to be built and somebody has to own it once it's running. Salesforce and HubSpot are easier, because managed connectors do most of the landing, and the real work becomes matching their customers to the ones in your ERP.

## What we look at

- Every system holding a number leadership asks about, and who owns each one
- Which sources have a managed connector, which need a pipeline built, and which only ever export files
- How each source reports change: a change feed, a modified timestamp, or nothing at all, which means full reads
- The request limits on the source side, so reporting loads don't spend the quota your users need during the day
- Which keys join the systems, the customer in the CRM against the same customer in the ERP, and what happens when they disagree
- Who may see what once it lands, written as Unity Catalog grants before the first table exists

## What you get

- Managed ingestion through Lakeflow Connect where it covers the source, such as Salesforce, Workday, Jira and SharePoint, HubSpot (CRM objects in Beta) and change data capture from SQL Server, MySQL and PostgreSQL
- Dynamics 365 apps through the Lakeflow Connect Dynamics 365 connector, which reads the Azure Synapse Link export in Data Lake Storage
- Business Central pipelines built by hand: incremental reads of the API pages filtered on lastModifiedDateTime, or the tables the bc2adls extension exports
- Bronze, silver and gold layers with the joins across systems done once, in one place, instead of again in every report
- Grants, row filters and column masks on the integrated tables, so ledger data stays with the people allowed to see it
- Freshness and row-count checks on each source, so a system that stops sending is noticed before a dashboard goes quiet

## Questions

### What are data integration services?

Getting data out of the systems a business runs on and into one place where it can be joined, governed and trusted. For us that place is Databricks. The work is deciding how each source is read, building the pipelines that keep it current, and resolving the records that don't match from one system to the next.

### Does Databricks have a connector for Business Central?

Not a managed one. Lakeflow Connect's Dynamics 365 connector reads data that Azure Synapse Link exports from Dataverse, and Business Central stores its data in its own database; only the tables you choose to sync ever reach Dataverse. Two routes work: a Databricks job reading the Business Central API pages incrementally on lastModifiedDateTime, or the bc2adls extension exporting tables to Azure Data Lake Storage for Databricks to read. Microsoft archived its own bc2adls repository in September 2023, and the community fork is the one still released.

### Can you connect Salesforce or HubSpot to Databricks?

Both have managed connectors in Lakeflow Connect, which turns most of the ingestion into setup. HubSpot's is the one to plan around. Its CRM objects are in Beta behind a workspace preview, HubSpot's API rate limits pace every sync, and tables the API can't read incrementally are re-read in full on each update. It also only authorises through a browser OAuth flow, so someone who can sign in to HubSpot has to be there when the connection is created. The real effort comes after landing: matching CRM accounts to ERP customers, deciding which CRM fields anyone trusts, and setting who can see pipeline and revenue.

### When is a managed connector the wrong choice?

When it doesn't cover the source, when you need a table or field it doesn't expose, or when the source only drops files. The Dynamics 365 connector, for instance, takes at most 250 tables per pipeline, a column you add to the selection later needs a manual full refresh before its history arrives, and schema evolution for CSV exports of standard entities is still in Private Preview. We read those limits for each source before we put a date on anything.

### What does an integration we've built look like?

The one we've written up is reporting for a manufacturing plant in its first year, built and tested on sample data shaped like the plant's own systems before any live system was connected. Business Central sits beside the scale house, the line historian, the lab and market prices: seven sources feeding 12 bronze, 15 silver and 10 gold tables, behind nine leadership dashboards and a Genie space. One report matches every Business Central receipt to its scale ticket, which neither system could do alone.

### Is this the same as connecting the systems to each other?

They're related jobs. Landing data in Databricks serves reporting, analytics and AI. Keeping operational systems in step, so an order placed on a website reaches the ERP without anyone re-keying it, is a separate job, and we do that too. For AmTab, Power Automate keeps the store, the CRM and Business Central holding the same orders, customers and invoices.

