In today’s video, we demonstrate how to use iceDQ to compare the data in a control file with the data in an actual file. The control file contains essential information about the actual data, such as record count, file size, and delimiter used. By comparing these two files, we can ensure their accuracy and identify any discrepancies.
We walk through setting up a checksum rule to compare the record count in the source file (customer.csv) with the record count in the control file (customer_stats.csv). This validation ensures that the data in the source file matches the expectations defined in the control file.
Key Highlights
*Checksum Validation: Learn how to compare data between a source file and a control file, using the record count and other metadata.
Control File Setup: Set up a control file to hold metadata like the expected record count and file size.
Data Accuracy: Validate that the source file data adheres to the specifications in the control file to ensure consistency and correctness.
Efficient Data Validation: Automate checks to quickly identify any discrepancies in record counts or other critical data metrics.
With iceDQ, you can automate this validation process, improve your data quality, and ensure data integrity across your ETL processes and data migrations.
Ready to automate your data validation?
Request a demo today and see how iceDQ can help you streamline your data testing and ensure the quality of your data.
00:00 - Introduction
00:08 - Control File vs Actual File Explained
00:26 - Creating a Checksum Rule
00:39 - Setting Up Source File Connection
01:25 - Setting Up Control File (Target) Connection
02:25 - Validating Record Count
03:05 - Publishing and Verifying the Rule
Request a Demo: https://icedq.com/request-a-demo
-------------------------------------------------
About iceDQ: Ensuring Reliable Data From Development to Production with iceDQ.
iceDQ is a one-stop platform for data reliability with unified data testing, monitoring, and observability. Large banks, insurance, healthcare, and other enterprises rely on iceDQ in both development and production environments, ensuring data reliability and robust processes.
Streamlined Data Testing in Development: iceDQ is used to automate data migration testing, ETL data pipeline testing, big data lake testing, BI report testing, and more. It helps identify and fix data issues early in the data development lifecycle.
Proactive Monitoring and Observability in Production: iceDQ is used by operations to establish checks and controls for their data pipelines, and the AI-based observability engine ensures anomalies are detected and incidents are reported.
-------------------------------------------------
Request a Demo: https://icedq.com/request-a-demo
Data Testing: https://icedq.com/product/data-testin...
Data Monitoring: https://icedq.com/product/data-monito...
Data Observability: https://icedq.com/product/data-observ...
Data Reliability: https://icedq.com/data-reliability-en...
LinkedIn: / icedq
Facebook: / icedq.toranainc
X: https://x.com/iceDQ_Toranainc
Reddit: / icedq
-------------------------------------------------
Don't forget to like this video, subscribe to our channel for more informative content, and hit the notification bell to stay updated with our latest uploads. Thank you for watching.
#iceDQ #ControlFileValidation #ChecksumValidation #ETLTesting #DataValidation #AutomatedTesting