Data Quality
166 subscribers
63 photos
3 videos
4 files
67 links
DQ about data qa
Download Telegram
We need to use data for functional and non-functional testing and of course for data quality. I'm sharing several sources which can help you to spawn data in the easiest way.

Faker is a Python library that generates fake data, such as names, addresses, and phone numbers. It combines predefined data with randomized values to produce a wide range of realistic and varied outputs and supports multiple languages and data formats. It is useful for testing software, populating databases with fake data, and creating fictional examples for presentations and documentation.

generatedata.com is a useful tool to quickly generate large quantities of fake data for a variety of purposes. It provides a simple interface for users to specify the type and quantity of data they want, as well as options for customizing the data, such as specifying data formats or choosing from predefined data sets. The website prepares the fake data in a variety of formats (CSV, JSON, SQL, Excel), and allows users to download the data or copy it to their clipboard.

Mackaro is pretty similar to generatedata.com. The main difference is that Mockaroo provides an API that allows users to generate fake data programmatically, making it easy to integrate data generation into users’ applications or scripts.

The Synthetic Data Vault (SDV) is a Synthetic Data Generation ecosystem of libraries that allows users to easily learn single-table, multi-table, and time-series datasets to generate new Synthetic Data afterward which has the same format and statistical properties as the original dataset.

Real DataSets:
Sometimes is important to use real data sets because they often provide more accurate and realistic results than fake data sets designed to be fictional.
Kaggle data sets:
https://www.kaggle.com/datasets
Mavenanalytics datasets:
https://www.mavenanalytics.io/data-playground
New York taxi trip records:
https://www.nyc.gov/site/tlc/about/tlc-trip-record-data.page
DataQA requires vast experience and knowledge as it combines two areas complicated by themselves. That is the reason behind the small count of Junior DataQA specialists. They have to learn a lot before entering data testing. Do you agree?
https://towardsdatascience.com/why-data-quality-is-harder-than-code-quality-a7ab78c9d9e
This week I was conducting tests for our #dataquality solution Data Quality Gate (DQG). As DGQ uses AWS, I had to find a way to emulate AWS locally. I'd like to share some tools I have found to help you test your solutions w/o AWS.
1. LocalStack is a fully functional local AWS cloud stack that emulates most AWS services. I used S3 and tried to use Lambda. LS has two versions: community with s3, ec2, and others; Pro with advanced services like EKS, SageMaker, etc. I didn’t find any issues with S3 but when using Lambda I noticed that the community version supports only zip archive deployment. Deployment via docker image is available only in the Pro version. Despite it, I strongly recommend this tool for local development.
2. Moto is a similar tool to Local Stack but has some limitations in services. However, Moto can be used as a mock in Python code that allows using it for unit testing. Interesting fact: Local stack uses Moto for covering services
3. MinIO is multi-cloud object storage. It doesn’t mock AWS but can be used as S3 storage. I didn’t try this tool but I heard good reviews from my colleagues.
In my experience, debugging pipelines with Jenkins was a real hassle. I would have to push any changes with every mistake and check remotely. I found the process incredibly frustrating. Nowadays, I work with GitHub Actions and create test pipelines for Data Quality Gate. Fortunately, there is a tool called Act that helps me to run my pipelines locally, without having to push code remotely. This saves me time and hassle.

The only downside of Act is that it doesn't support service containers. Link to issue

If you're working with GitLab CI, there is a similar tool called gitlab-ci-local

Did you try similar tools?
True story
Learn how to read data from a specific sheet in an Excel file using Great Expectations! My tip for today shows you how to identify the sheet name in runtime_parameters to streamline your data reading process. Check out the code snippet below to get started:
Great Expectations provides a special structure for projects, as described in the official documentation. You can create this structure through the official cli, but this requires a certain amount of user interactions. This can be problematic if you need to set up a GX project in a Docker container or AWS Lambda.

For this purpose, you can use the following code to initiate a context in a particular directory
As data quality professionals, we often need to print logs for debugging purposes. Luckily, setting up a basic configuration for logging in to Great Expectations is simple! Check out the example to get started:
If you're using SQLAlchemy and a database with Great Expectations, here's a helpful tip: You can use logging to view the queries that GX executes for validations. By enabling logging, you can gain insights into the database interactions and improve your understanding of how GX works with your data.
If you want to speed up your validations in Great Expectations, try running them in parallel. One way to do this is using #pytest, which allows you to run tests concurrently. By wrapping your checkpoints in tests and running them in parallel, you can significantly reduce the time it takes to validate your data. So give it a try and see if it improves your workflow!
#DataQuality #tips #datamanagement #GXtips #data
Why is Data Quality Important?

I have one answer which should cover all doubts about DQ's importance. Agree?
👏3
Data Management as Real Estate Management. Does it work for you?
Source
Friday !
The first course for GX. You can try GX in action without struggle with setting: Great Expectations, a data validation library for Python
True story.
Hey, do you use AWS Athena to analyze your data? If so, you might want to check out this awesome article. It shows you two ways to keep your data quality in check: AWS Glue Data Quality and Great Expectations. Looks like GX is still the best tool for Data QA
👍1
📢 Good news from Soda🎉

This week, Soda released Rule Suggestions, joining the ranks of GreatExpectation and Deequ in offering this powerful feature. 🚀 Now, you have multiple options to generate initial tests for your data source and improve your data quality management. 💪

Soda currently supports 6 rules for generating these suggestions, and you can find more information in the official documentation. 📋 Make the most out of these amazing tools and optimize your data with ease. 🌟
#DataQuality #SodaUpdates #GreatExpectation #Deequ #RuleSuggestions #Efficiency