Data Quality
166 subscribers
63 photos
3 videos
4 files
67 links
DQ about data qa
Download Telegram
Channel created
This media is not supported in your browser
VIEW IN TELEGRAM
We need to use data for functional and non-functional testing and of course for data quality. I'm sharing several sources which can help you to spawn data in the easiest way.

Faker is a Python library that generates fake data, such as names, addresses, and phone numbers. It combines predefined data with randomized values to produce a wide range of realistic and varied outputs and supports multiple languages and data formats. It is useful for testing software, populating databases with fake data, and creating fictional examples for presentations and documentation.

generatedata.com is a useful tool to quickly generate large quantities of fake data for a variety of purposes. It provides a simple interface for users to specify the type and quantity of data they want, as well as options for customizing the data, such as specifying data formats or choosing from predefined data sets. The website prepares the fake data in a variety of formats (CSV, JSON, SQL, Excel), and allows users to download the data or copy it to their clipboard.

Mackaro is pretty similar to generatedata.com. The main difference is that Mockaroo provides an API that allows users to generate fake data programmatically, making it easy to integrate data generation into users’ applications or scripts.

The Synthetic Data Vault (SDV) is a Synthetic Data Generation ecosystem of libraries that allows users to easily learn single-table, multi-table, and time-series datasets to generate new Synthetic Data afterward which has the same format and statistical properties as the original dataset.

Real DataSets:
Sometimes is important to use real data sets because they often provide more accurate and realistic results than fake data sets designed to be fictional.
Kaggle data sets:
https://www.kaggle.com/datasets
Mavenanalytics datasets:
https://www.mavenanalytics.io/data-playground
New York taxi trip records:
https://www.nyc.gov/site/tlc/about/tlc-trip-record-data.page
DataQA requires vast experience and knowledge as it combines two areas complicated by themselves. That is the reason behind the small count of Junior DataQA specialists. They have to learn a lot before entering data testing. Do you agree?
https://towardsdatascience.com/why-data-quality-is-harder-than-code-quality-a7ab78c9d9e
This week I was conducting tests for our #dataquality solution Data Quality Gate (DQG). As DGQ uses AWS, I had to find a way to emulate AWS locally. I'd like to share some tools I have found to help you test your solutions w/o AWS.
1. LocalStack is a fully functional local AWS cloud stack that emulates most AWS services. I used S3 and tried to use Lambda. LS has two versions: community with s3, ec2, and others; Pro with advanced services like EKS, SageMaker, etc. I didn’t find any issues with S3 but when using Lambda I noticed that the community version supports only zip archive deployment. Deployment via docker image is available only in the Pro version. Despite it, I strongly recommend this tool for local development.
2. Moto is a similar tool to Local Stack but has some limitations in services. However, Moto can be used as a mock in Python code that allows using it for unit testing. Interesting fact: Local stack uses Moto for covering services
3. MinIO is multi-cloud object storage. It doesn’t mock AWS but can be used as S3 storage. I didn’t try this tool but I heard good reviews from my colleagues.
In my experience, debugging pipelines with Jenkins was a real hassle. I would have to push any changes with every mistake and check remotely. I found the process incredibly frustrating. Nowadays, I work with GitHub Actions and create test pipelines for Data Quality Gate. Fortunately, there is a tool called Act that helps me to run my pipelines locally, without having to push code remotely. This saves me time and hassle.

The only downside of Act is that it doesn't support service containers. Link to issue

If you're working with GitLab CI, there is a similar tool called gitlab-ci-local

Did you try similar tools?
True story