Load data from Google Cloud Storage [Managed Deployment]
Load data into your Managed Simon deployment directly from files in your Google Cloud Storage bucket
Overview
Simon can load data directly from files in your Google Cloud Storage (GCS) bucket. Our Snowflake account reads your files in bulk on a schedule, so for large tables this is often faster and cheaper than extracting from a database. Here's how to get set up.
For Managed deploymentsFile loading is for Managed deployments, which run on a Simon-owned Snowflake account. In a Connected deployment, Simon reads directly from your own Snowflake account instead.
Set up the connection 🔐
- Export the data you want to send to a GCS bucket you own. Give each dataset its own path. Supported formats are Parquet (recommended), JSON, CSV, and TSV, with optional compression.
- Send us the bucket name and path. We'll reply with the Google service account Simon uses to read your files.
- In the Google Cloud console, create a custom role with the permissions below.
- Go to Cloud Storage > Buckets and select your bucket. On the Permissions tab, click Grant access, paste the service account we sent you, assign your custom role, and click Save.
| Permission | Required | What it allows |
|---|---|---|
storage.buckets.get | Yes | Read the bucket's details, including its location |
storage.objects.get | Yes | Read your files |
storage.objects.list | Yes | Find new files in your path |
storage.objects.create | No | Write files to the bucket |
storage.objects.delete | No | Remove files from the bucket |
Send us the dataset details
For each dataset, send us the following.
| Detail | Description |
|---|---|
| Bucket path | Where the dataset's files live |
| File format | Parquet, JSON, CSV, or TSV, and any compression |
| Columns | The columns you're sending and the data type of each |
| Record ID column | A column that uniquely identifies each record. Simon uses it to match records, so each value must appear only once |
| Updated timestamp column | The column that records when each record last changed. Required for incremental loads |
| Refresh schedule | How often Simon should refresh the dataset, and at what time |
Choose a load pattern 🔄
Incremental
Incremental loads work best for large or growing tables, like orders or email events.
- Write only new or changed records each run, into date-separated folders:
<your path>/YYYY/MM/DD/ - Simon loads each new file once and merges it into the existing table using your record ID column
- To delete records, write files containing their record IDs to
<your path>/deletes/- Use the same file format as your data files
- The record ID column is the only required column
- Simon applies deletes after it merges new data on each run
Overwrite
Overwrite loads work best for smaller reference or snapshot tables, like a product catalog or store list.
- Write a full copy of the table each run, into a new dated folder:
<your path>/YYYY/MM/DD/ - If you don't want to keep historical files, write each copy to the same folder instead
- Simon always reads from the most recent dated folder
- Simon reads the files in place as a Snowflake external table, so each query uses the latest files
Data egress feesYour cloud provider may charge data egress fees when Simon reads files from your bucket, especially across regions or clouds. Check with your Google Cloud account team about pricing for your bucket's location.
What's next
Once your first files load, configure the new tables in the Schema Builder so your team can use them in segments and content.
Updated about 5 hours ago
