Load data from Azure Blob Storage [Managed Deployment]

Load data into your Managed Simon deployment directly from files in your Azure Blob Storage account

Overview

Simon can load data directly from files in your Azure Blob Storage account. Our Snowflake account reads your files in bulk on a schedule, so for large tables this is often faster and cheaper than extracting from a database. Here's how to get set up.

📘

For Managed deployments

File loading is for Managed deployments, which run on a Simon-owned Snowflake account. In a Connected deployment, Simon reads directly from your own Snowflake account instead.

Set up the connection 🔐

  1. Export the data you want to send to your storage account. Give each dataset its own path. Supported formats are Parquet (recommended), JSON, CSV, and TSV, with optional compression.
  2. Send us the following so we can create the connection.
DetailDescription
Tenant IDYour Azure tenant ID
RegionThe Azure region your storage account is in
PathThe storage account, container, and path where your files will live
  1. We'll reply with two values you'll use to grant Simon access.
ValueDescription
Consent URLA Microsoft link that lets Simon's Snowflake account request access to your tenant
App nameThe name of the Snowflake service principal you'll grant access to

Grant Simon access to your storage account

  1. Open the consent URL in a web browser. On the Microsoft permissions request page, click Accept. This lets the Snowflake service principal request an access token in your tenant. The token only works after you assign it a role in the steps below.
  2. Sign in to the Microsoft Azure portal.
  3. Go to Azure Services > Storage Accounts and click the storage account that holds your files.
  4. Click Access Control (IAM) > Add role assignment.
  5. Select the Storage Blob Data Reader role. It grants read-only access, which is all Simon needs to load your files.
  6. On the Members tab, click Select members and search for the app name we sent you. Search only for the part before the underscore.
  7. Click Review + assign.
📘

The service principal can take time to appear

After you accept the consent request, Azure can take an hour or longer to create the Snowflake service principal. If your search in step 6 finds nothing, wait and try again. Role assignments can take up to five minutes to take effect.

Send us the dataset details

For each dataset, send us the following.

DetailDescription
File formatParquet, JSON, CSV, or TSV, and any compression
ColumnsThe columns you're sending and the data type of each
Record ID columnA column that uniquely identifies each record. Simon uses it to match records, so each value must appear only once
Updated timestamp columnThe column that records when each record last changed. Required for incremental loads
Refresh scheduleHow often Simon should refresh the dataset, and at what time

Choose a load pattern 🔄

Incremental

Incremental loads work best for large or growing tables, like orders or email events.

  • Write only new or changed records each run, into date-separated folders: <your path>/YYYY/MM/DD/
  • Simon loads each new file once and merges it into the existing table using your record ID column
  • To delete records, write files containing their record IDs to <your path>/deletes/
    • Use the same file format as your data files
    • The record ID column is the only required column
  • Simon applies deletes after it merges new data on each run

Overwrite

Overwrite loads work best for smaller reference or snapshot tables, like a product catalog or store list.

  • Write a full copy of the table each run, into a new dated folder: <your path>/YYYY/MM/DD/
  • If you don't want to keep historical files, write each copy to the same folder instead
  • Simon always reads from the most recent dated folder
  • Simon reads the files in place as a Snowflake external table, so each query uses the latest files
🚧

Data egress fees

Your cloud provider may charge data egress fees when Simon reads files from your storage account, especially across regions or clouds. Check with your Azure account team about pricing for your storage account's location.

What's next

Once your first files load, configure the new tables in the Schema Builder so your team can use them in segments and content.


Did this page help you?