Migrating from a legacy data warehouse to Snowflake requires validation at each step. This guide outlines a four-step process using Datagaps DataFlow: comparing extracted data (row counts, encoding, completeness) between the legacy warehouse and AWS S3, validating transformations using Data Profile and Data Rules components, verifying data loaded into Snowflake against staging files, and end-to-end validation. It also covers using BI Validator to test report accuracy, layout, performance, and security after switching reports to Snowflake.
Key Takeaways
- Extraction validation confirms lossless data movement — comparing row counts, data encoding, completeness, and values between the legacy warehouse and AWS S3 landing zone ensures nothing was lost during extraction.
- Transformation step requires data quality checks before load — Data Profile and Data Rules components help identify and curate data quality issues before syncing data to the staging zone, ahead of loading into Snowflake.
- DataFlow supports single-test, end-to-end validation — a single DataFlow can validate data between the legacy warehouse and S3, as well as between the legacy warehouse and Snowflake, in one unified test.
- Reports need separate validation after migration — BI Validator checks report data accuracy, layout, performance, security, and stress-tests reports under concurrent user loads to confirm they function correctly once switched to Snowflake.
Migrating to Snowflake Cloud Data Warehouse
Snowflake data migration testing is the process of validating data at each stage of a move from a legacy data warehouse — such as Netezza — to a cloud-based Snowflake data warehouse: extraction, transformation, load, and the reports built on top of it. Datagaps DataFlow can be used to validate data at each of these steps, as well as run end-to-end validation across the entire migration. If you are looking for Snowflake testing tools – Try Dataflow free for 14 days

| Step | What Gets Validated |
|---|---|
| 1. Extract from legacy warehouse | Row counts, data encoding, completeness, and values between the legacy warehouse and the AWS S3 landing zone |
| 2. Transform data | Data quality issues via profiling and data rules; comparison between landing and staging (curated) zones |
| 3. Copy data to Snowflake | Row counts, encoding, completeness, and values between S3 staging files and Snowflake; end-to-end legacy-to-Snowflake validation |
| 4. Modify reports to use Snowflake | Report data, layout, performance, security, and behavior under concurrent user load, compared against the legacy reports |
Step 1: Extract data from the Legacy data warehouse
Data is typically extracted into CSV or Parquet format and moved to a landing zone in AWS S3. Depending on the data volumes, AWS offers multiple options for moving the files to S3. Once the data has been moved to AWS S3, data validations need to be performed to ensure that all the data was properly extracted and migrated to AWS S3. Since there are not many transformations in this step, these tests are typically one-to-one comparisons of the data in the tables in the legacy data warehouse and the files in the AWS S3 landing zone.
– Compare table to file row counts
– Compare data encoding
– Compare data completeness
– Compare data values

Data comparison test case

The output of data comparison
Step 2: Transform data
Transformations such as data type conversions can be performed in this step. Data curation can be also done to improve the Data Quality before the data is loaded into Snowflake. Before curating the data, it is important to profile the data and run data quality tests to identify data quality issues with the data. DataFlow can be used to perform these tasks.
– Compare data between landing zone and staging (curated) zone in S3
– Use Data Profile and Data Rules components to identify data quality issues
– Curate data and sync to the staging zone


Data Rules component
Step 3: Copy data to Snowflake
– Compare table to file row counts
– Compare data encoding
– Compare data completeness
– Compare data values
– End-to-end data validation (Legacy data warehouse to Snowflake)
DataFlow can be used to perform end-to-end Data Validation in a single test as shown to the right. A single DataFlow can be used to compare data between legacy data warehouse and S3 as well as legacy data warehouse and Snowflake.

End-to-end test case
Step 4: Modify reports to Use Snowflake
While snowflake provides JDBC/ODBC drivers and supports most of the commonly used SQL functions, there are going to be some differences between the way reports are developed and executed in the legacy Data Warehouse and Snowflake. Once these changes are made, thorough testing needs to be performed between the reports using the legacy data warehouse and the equivalent reports using Snowflake.
– Compare report data
– Compare report layout
– Compare report performance
– Stress test reports in the new environments by simulating concurrent user loads
– Compare security
Datagaps BI Validator is a no-code BI Testing Tool that can help automate all these tests for the supported BI tools.
Try BI Validator free for 14 days for your Snowflake BI testing needs – Download Now
Conclusion
Migrating to Snowflake is never just a lift-and-shift — it’s a multi-stage process where errors introduced at extraction or transformation can silently propagate all the way through to the reports business users rely on every day. By validating data at each stage — comparing row counts and values as data moves from the legacy warehouse to S3, profiling and curating data during transformation, and reconciling loaded data against Snowflake — teams can catch discrepancies before they compound. And because reports rarely translate one-to-one across platforms, testing report accuracy, layout, performance, and security after the switch is just as critical as validating the underlying data. With tools like Datagaps DataFlow and BI Validator automating these checks at every step, organizations can move to Snowflake with confidence that their data — and the reports built on it — remain accurate and trustworthy from day one.
FAQs: Snowflake Data Migration Validation
1) What are the main steps involved in validating a Snowflake data migration?
Snowflake migration validation typically includes four stages: validating data extraction from the legacy data warehouse to AWS S3, verifying data transformations and quality in the staging area, comparing staged data with data loaded into Snowflake, and validating BI reports after they are connected to the new Snowflake environment.
2) How is data extraction validated when migrating to Snowflake?
Extraction validation compares row counts, data completeness, encoding, and actual data values between the legacy source system and the files written to AWS S3. This ensures data is transferred accurately before any transformation or loading takes place.
3) Why is data profiling important before loading data into Snowflake?
Data profiling identifies quality issues such as missing values, invalid formats, and inconsistencies before data is loaded into Snowflake. Cleaning and validating data in the staging area helps ensure only accurate, high-quality data reaches the target warehouse.
4) What does report validation involve after migrating to Snowflake?
Post-migration report validation verifies that dashboards and reports produce the same results as before migration. It includes checking data accuracy, visual consistency, report performance, stress testing under concurrent usage, and confirming that security and user access permissions remain correctly configured.

Rajesh Kumar A
Digital Marketing Manager, Datagaps
Digital Marketing Manager at Datagaps. Drives data-driven growth through content, performance campaigns, and marketing technology.

Subrahmanya Narayana Chirravuri
Senior Director, Technology, Datagaps
Senior Director of Technology at Datagaps. Leads engineering for the ETL, BI, and data-quality validation platforms.





