Unify and manage your data

Best practices for setting up Reltio data sharing with Databricks - essentials

Follow these practices when you set up a Databricks data share so that the initial load and the ongoing sync perform as expected.

Complete the initial data load before you set up data share

Complete the initial data load into Reltio, and allow all match & merge operations to complete before you set up the Databricks data share.

This prevents unintended events from being synced and protects overall sync performance.

Disable data share during initial load (if already enabled)

If you set up data share before completing the initial data load, disable it until all the initial data load and all the related match and merge operations have completed and you have verified that the data unification outcomes match your expectations.

To disable data share, follow the below steps.

  1. Obtain the current physical configuration of your tenant using the below endpoint.

    GET {Env_URL}/reltio/tenants/{TenantId}/dataPipelineConfig
  2. In the JSON response, locate the desired data share within the adapters array, set its enabled parameter to false, and save the configuration.

    "adapters": [
      {
        "name": "<...>", // Find the exact data share that you want to disable
        "enabled": false,
        ...
      }
      ...
    ]
  3. Post the updated physical configuration to the tenant using the below endpoint.
    PUT {ENVIRONMENT_URL}/reltio/tenants/{TenantId}/dataPipelineConfig

    'Body' parameter should be set to JSON with the content of updated dataPipelineConfig.

Setup a new or enable an existing data share

After the initial full data load is complete, verify the following:

  • All related match and merge operations have completed.

  • Data unification outcomes meet your expectations.

To enable an existing data share, update the enabled parameter for the required data share to true by following below steps.

  1. Obtain the current physical configuration of your tenant using the below endpoint.

    GET {Env_URL}/reltio/tenants/{TenantId}/dataPipelineConfig
  2. In the JSON response, locate the desired data share within the adapters array, set its enabled parameter to true, and save the configuration.

    "adapters": [
      {
        "name": "<...>", // Find the exact data share that you want to enable
        "enabled": true,
        ...
      }
      ...
    ]
  3. Post the updated physical configuration to the tenant using the below endpoint.
    PUT {ENVIRONMENT_URL}/reltio/tenants/{TenantId}/dataPipelineConfig
    'Body' parameter should be set to JSON with the content of updated dataPipelineConfig.

Run the one time activity of initial full data sync in sequence

Once the data share setup is complete and it is enabled, all subsequent events in Reltio, post the data share setup are automatically synchronized across all supported data objects. However, to initiate the full sync of the data that existed before the data share setup, trigger the syncToDataPipeline API once for each data type in the given sequence.

  1. Sync all entity data sets by using the syncToDataPipeline API with the dataTypes parameter set to entities.

  2. Sync all relation data sets by using the syncToDataPipeline API with the dataTypes parameter set to relations.

  3. Sync matches, merges, and links data sets by using the syncToDataPipeline API for each data type respectively.

  4. Sync any other data sets that you want, such as interactions, using the syncToDataPipeline API with the dataTypes parameter set to the required data type.

    For more information on how to sync the data, see API Guide.

Practices to avoid

Avoid enabling data share before initial data load

Do not enable data share until the initial load is complete and the data is validated. This prevents unnecessary or unintended events from being synced and protects overall sync performance.

Avoid syncing all data sets at once

Do not trigger a full sync for all data sets simultaneously, as this can negatively impact overall sync performance.