Load data from Google Cloud Storage [Managed Deployment]

Load data into your Managed Simon deployment directly from files in your Google Cloud Storage bucket

Overview

Simon can load data directly from files in your Google Cloud Storage (GCS) bucket. Our Snowflake account reads your files in bulk on a schedule, so for large tables this is often faster and cheaper than extracting from a database. Here's how to get set up.

📘

For Managed deployments

File loading is for Managed deployments, which run on a Simon-owned Snowflake account. In a Connected deployment, Simon reads directly from your own Snowflake account instead.

Set up the connection 🔐

  1. Export the data you want to send to a GCS bucket you own. Give each dataset its own path. Supported formats are Parquet (recommended), JSON, CSV, and TSV, with optional compression.
  2. Send us the bucket name and path. We'll reply with the Google service account Simon uses to read your files.
  3. In the Google Cloud console, create a custom role with the permissions below.
  4. Go to Cloud Storage > Buckets and select your bucket. On the Permissions tab, click Grant access, paste the service account we sent you, assign your custom role, and click Save.
PermissionRequiredWhat it allows
storage.buckets.getYesRead the bucket's details, including its location
storage.objects.getYesRead your files
storage.objects.listYesFind new files in your path
storage.objects.createNoWrite files to the bucket
storage.objects.deleteNoRemove files from the bucket

Send us the dataset details

For each dataset, send us the following.

DetailDescription
Bucket pathWhere the dataset's files live
File formatParquet, JSON, CSV, or TSV, and any compression
ColumnsThe columns you're sending and the data type of each
Record ID columnA column that uniquely identifies each record. Simon uses it to match records, so each value must appear only once
Updated timestamp columnThe column that records when each record last changed. Required for incremental loads
Refresh scheduleHow often Simon should refresh the dataset, and at what time

Choose a load pattern 🔄

Incremental

Incremental loads work best for large or growing tables, like orders or email events.

  • Write only new or changed records each run, into date-separated folders: <your path>/YYYY/MM/DD/
  • Simon loads each new file once and merges it into the existing table using your record ID column
  • To delete records, write files containing their record IDs to <your path>/deletes/
    • Use the same file format as your data files
    • The record ID column is the only required column
  • Simon applies deletes after it merges new data on each run

Overwrite

Overwrite loads work best for smaller reference or snapshot tables, like a product catalog or store list.

  • Write a full copy of the table each run, into a new dated folder: <your path>/YYYY/MM/DD/
  • If you don't want to keep historical files, write each copy to the same folder instead
  • Simon always reads from the most recent dated folder
  • Simon reads the files in place as a Snowflake external table, so each query uses the latest files
🚧

Data egress fees

Your cloud provider may charge data egress fees when Simon reads files from your bucket, especially across regions or clouds. Check with your Google Cloud account team about pricing for your bucket's location.

What's next

Once your first files load, configure the new tables in the Schema Builder so your team can use them in segments and content.


Did this page help you?