Load data from Azure Blob Storage [Managed Deployment]
Load data into your Managed Simon deployment directly from files in your Azure Blob Storage account
Overview
Simon can load data directly from files in your Azure Blob Storage account. Our Snowflake account reads your files in bulk on a schedule, so for large tables this is often faster and cheaper than extracting from a database. Here's how to get set up.
For Managed deploymentsFile loading is for Managed deployments, which run on a Simon-owned Snowflake account. In a Connected deployment, Simon reads directly from your own Snowflake account instead.
Set up the connection 🔐
- Export the data you want to send to your storage account. Give each dataset its own path. Supported formats are Parquet (recommended), JSON, CSV, and TSV, with optional compression.
- Send us the following so we can create the connection.
| Detail | Description |
|---|---|
| Tenant ID | Your Azure tenant ID |
| Region | The Azure region your storage account is in |
| Path | The storage account, container, and path where your files will live |
- We'll reply with two values you'll use to grant Simon access.
| Value | Description |
|---|---|
| Consent URL | A Microsoft link that lets Simon's Snowflake account request access to your tenant |
| App name | The name of the Snowflake service principal you'll grant access to |
Grant Simon access to your storage account
- Open the consent URL in a web browser. On the Microsoft permissions request page, click Accept. This lets the Snowflake service principal request an access token in your tenant. The token only works after you assign it a role in the steps below.
- Sign in to the Microsoft Azure portal.
- Go to Azure Services > Storage Accounts and click the storage account that holds your files.
- Click Access Control (IAM) > Add role assignment.
- Select the Storage Blob Data Reader role. It grants read-only access, which is all Simon needs to load your files.
- On the Members tab, click Select members and search for the app name we sent you. Search only for the part before the underscore.
- Click Review + assign.
The service principal can take time to appearAfter you accept the consent request, Azure can take an hour or longer to create the Snowflake service principal. If your search in step 6 finds nothing, wait and try again. Role assignments can take up to five minutes to take effect.
Send us the dataset details
For each dataset, send us the following.
| Detail | Description |
|---|---|
| File format | Parquet, JSON, CSV, or TSV, and any compression |
| Columns | The columns you're sending and the data type of each |
| Record ID column | A column that uniquely identifies each record. Simon uses it to match records, so each value must appear only once |
| Updated timestamp column | The column that records when each record last changed. Required for incremental loads |
| Refresh schedule | How often Simon should refresh the dataset, and at what time |
Choose a load pattern 🔄
Incremental
Incremental loads work best for large or growing tables, like orders or email events.
- Write only new or changed records each run, into date-separated folders:
<your path>/YYYY/MM/DD/ - Simon loads each new file once and merges it into the existing table using your record ID column
- To delete records, write files containing their record IDs to
<your path>/deletes/- Use the same file format as your data files
- The record ID column is the only required column
- Simon applies deletes after it merges new data on each run
Overwrite
Overwrite loads work best for smaller reference or snapshot tables, like a product catalog or store list.
- Write a full copy of the table each run, into a new dated folder:
<your path>/YYYY/MM/DD/ - If you don't want to keep historical files, write each copy to the same folder instead
- Simon always reads from the most recent dated folder
- Simon reads the files in place as a Snowflake external table, so each query uses the latest files
Data egress feesYour cloud provider may charge data egress fees when Simon reads files from your storage account, especially across regions or clouds. Check with your Azure account team about pricing for your storage account's location.
What's next
Once your first files load, configure the new tables in the Schema Builder so your team can use them in segments and content.
Updated about 5 hours ago
