We need to use data for functional and non-functional testing and of course for data quality. I'm sharing several sources which can help you to spawn data in the easiest way.
Faker is a Python library that generates fake data, such as names, addresses, and phone numbers. It combines predefined data with randomized values to produce a wide range of realistic and varied outputs and supports multiple languages and data formats. It is useful for testing software, populating databases with fake data, and creating fictional examples for presentations and documentation.
generatedata.com is a useful tool to quickly generate large quantities of fake data for a variety of purposes. It provides a simple interface for users to specify the type and quantity of data they want, as well as options for customizing the data, such as specifying data formats or choosing from predefined data sets. The website prepares the fake data in a variety of formats (CSV, JSON, SQL, Excel), and allows users to download the data or copy it to their clipboard.
Mackaro is pretty similar to generatedata.com. The main difference is that Mockaroo provides an API that allows users to generate fake data programmatically, making it easy to integrate data generation into users’ applications or scripts.
The Synthetic Data Vault (SDV) is a Synthetic Data Generation ecosystem of libraries that allows users to easily learn single-table, multi-table, and time-series datasets to generate new Synthetic Data afterward which has the same format and statistical properties as the original dataset.
Real DataSets:
Sometimes is important to use real data sets because they often provide more accurate and realistic results than fake data sets designed to be fictional.
Kaggle data sets:
https://www.kaggle.com/datasets
Mavenanalytics datasets:
https://www.mavenanalytics.io/data-playground
New York taxi trip records:
https://www.nyc.gov/site/tlc/about/tlc-trip-record-data.page
Faker is a Python library that generates fake data, such as names, addresses, and phone numbers. It combines predefined data with randomized values to produce a wide range of realistic and varied outputs and supports multiple languages and data formats. It is useful for testing software, populating databases with fake data, and creating fictional examples for presentations and documentation.
generatedata.com is a useful tool to quickly generate large quantities of fake data for a variety of purposes. It provides a simple interface for users to specify the type and quantity of data they want, as well as options for customizing the data, such as specifying data formats or choosing from predefined data sets. The website prepares the fake data in a variety of formats (CSV, JSON, SQL, Excel), and allows users to download the data or copy it to their clipboard.
Mackaro is pretty similar to generatedata.com. The main difference is that Mockaroo provides an API that allows users to generate fake data programmatically, making it easy to integrate data generation into users’ applications or scripts.
The Synthetic Data Vault (SDV) is a Synthetic Data Generation ecosystem of libraries that allows users to easily learn single-table, multi-table, and time-series datasets to generate new Synthetic Data afterward which has the same format and statistical properties as the original dataset.
Real DataSets:
Sometimes is important to use real data sets because they often provide more accurate and realistic results than fake data sets designed to be fictional.
Kaggle data sets:
https://www.kaggle.com/datasets
Mavenanalytics datasets:
https://www.mavenanalytics.io/data-playground
New York taxi trip records:
https://www.nyc.gov/site/tlc/about/tlc-trip-record-data.page
DataQA requires vast experience and knowledge as it combines two areas complicated by themselves. That is the reason behind the small count of Junior DataQA specialists. They have to learn a lot before entering data testing. Do you agree?
https://towardsdatascience.com/why-data-quality-is-harder-than-code-quality-a7ab78c9d9e
https://towardsdatascience.com/why-data-quality-is-harder-than-code-quality-a7ab78c9d9e
This week I was conducting tests for our #dataquality solution Data Quality Gate (DQG). As DGQ uses AWS, I had to find a way to emulate AWS locally. I'd like to share some tools I have found to help you test your solutions w/o AWS.
1. LocalStack is a fully functional local AWS cloud stack that emulates most AWS services. I used S3 and tried to use Lambda. LS has two versions: community with s3, ec2, and others; Pro with advanced services like EKS, SageMaker, etc. I didn’t find any issues with S3 but when using Lambda I noticed that the community version supports only zip archive deployment. Deployment via docker image is available only in the Pro version. Despite it, I strongly recommend this tool for local development.
2. Moto is a similar tool to Local Stack but has some limitations in services. However, Moto can be used as a mock in Python code that allows using it for unit testing. Interesting fact: Local stack uses Moto for covering services
3. MinIO is multi-cloud object storage. It doesn’t mock AWS but can be used as S3 storage. I didn’t try this tool but I heard good reviews from my colleagues.
1. LocalStack is a fully functional local AWS cloud stack that emulates most AWS services. I used S3 and tried to use Lambda. LS has two versions: community with s3, ec2, and others; Pro with advanced services like EKS, SageMaker, etc. I didn’t find any issues with S3 but when using Lambda I noticed that the community version supports only zip archive deployment. Deployment via docker image is available only in the Pro version. Despite it, I strongly recommend this tool for local development.
2. Moto is a similar tool to Local Stack but has some limitations in services. However, Moto can be used as a mock in Python code that allows using it for unit testing. Interesting fact: Local stack uses Moto for covering services
3. MinIO is multi-cloud object storage. It doesn’t mock AWS but can be used as S3 storage. I didn’t try this tool but I heard good reviews from my colleagues.
In my experience, debugging pipelines with Jenkins was a real hassle. I would have to push any changes with every mistake and check remotely. I found the process incredibly frustrating. Nowadays, I work with GitHub Actions and create test pipelines for Data Quality Gate. Fortunately, there is a tool called Act that helps me to run my pipelines locally, without having to push code remotely. This saves me time and hassle.
The only downside of Act is that it doesn't support service containers. Link to issue
If you're working with GitLab CI, there is a similar tool called gitlab-ci-local
Did you try similar tools?
The only downside of Act is that it doesn't support service containers. Link to issue
If you're working with GitLab CI, there is a similar tool called gitlab-ci-local
Did you try similar tools?
Great Expectations provides a special structure for projects, as described in the official documentation. You can create this structure through the official cli, but this requires a certain amount of user interactions. This can be problematic if you need to set up a GX project in a Docker container or AWS Lambda.
For this purpose, you can use the following code to initiate a context in a particular directory
For this purpose, you can use the following code to initiate a context in a particular directory
The article shows how you can implement different data quality dimensions with Great Expectations. It is an important topic because Data QA s have no standard here.
Please share your feedback
Link
Please share your feedback
Link
KDnuggets
Data Quality Dimensions: Assuring Your Data Quality with Great Expectations
This article highlights the significance of ensuring high-quality data and presents six key dimensions for measuring it. These dimensions include Completeness, Consistency, Integrity, Timelessness, Uniqueness, and Validity.
If you want to speed up your validations in Great Expectations, try running them in parallel. One way to do this is using #pytest, which allows you to run tests concurrently. By wrapping your checkpoints in tests and running them in parallel, you can significantly reduce the time it takes to validate your data. So give it a try and see if it improves your workflow!
#DataQuality #tips #datamanagement #GXtips #data
#DataQuality #tips #datamanagement #GXtips #data
The first course for GX. You can try GX in action without struggle with setting: Great Expectations, a data validation library for Python
Hey, do you use AWS Athena to analyze your data? If so, you might want to check out this awesome article. It shows you two ways to keep your data quality in check: AWS Glue Data Quality and Great Expectations. Looks like GX is still the best tool for Data QA
👍1
📢 Good news from Soda🎉
This week, Soda released Rule Suggestions, joining the ranks of GreatExpectation and Deequ in offering this powerful feature. 🚀 Now, you have multiple options to generate initial tests for your data source and improve your data quality management. 💪
Soda currently supports 6 rules for generating these suggestions, and you can find more information in the official documentation. 📋 Make the most out of these amazing tools and optimize your data with ease. 🌟
#DataQuality #SodaUpdates #GreatExpectation #Deequ #RuleSuggestions #Efficiency
This week, Soda released Rule Suggestions, joining the ranks of GreatExpectation and Deequ in offering this powerful feature. 🚀 Now, you have multiple options to generate initial tests for your data source and improve your data quality management. 💪
Soda currently supports 6 rules for generating these suggestions, and you can find more information in the official documentation. 📋 Make the most out of these amazing tools and optimize your data with ease. 🌟
#DataQuality #SodaUpdates #GreatExpectation #Deequ #RuleSuggestions #Efficiency
Soda Documentation
Adopt check Suggestions
The Check Suggestions CLI assisstant is designed to simplify the process of auto-generating basic data quality checks in SodaCL.