Data Quality
166 subscribers
63 photos
3 videos
4 files
67 links
DQ about data qa
Download Telegram
DataQA requires vast experience and knowledge as it combines two areas complicated by themselves. That is the reason behind the small count of Junior DataQA specialists. They have to learn a lot before entering data testing. Do you agree?
https://towardsdatascience.com/why-data-quality-is-harder-than-code-quality-a7ab78c9d9e
This week I was conducting tests for our #dataquality solution Data Quality Gate (DQG). As DGQ uses AWS, I had to find a way to emulate AWS locally. I'd like to share some tools I have found to help you test your solutions w/o AWS.
1. LocalStack is a fully functional local AWS cloud stack that emulates most AWS services. I used S3 and tried to use Lambda. LS has two versions: community with s3, ec2, and others; Pro with advanced services like EKS, SageMaker, etc. I didn’t find any issues with S3 but when using Lambda I noticed that the community version supports only zip archive deployment. Deployment via docker image is available only in the Pro version. Despite it, I strongly recommend this tool for local development.
2. Moto is a similar tool to Local Stack but has some limitations in services. However, Moto can be used as a mock in Python code that allows using it for unit testing. Interesting fact: Local stack uses Moto for covering services
3. MinIO is multi-cloud object storage. It doesn’t mock AWS but can be used as S3 storage. I didn’t try this tool but I heard good reviews from my colleagues.
In my experience, debugging pipelines with Jenkins was a real hassle. I would have to push any changes with every mistake and check remotely. I found the process incredibly frustrating. Nowadays, I work with GitHub Actions and create test pipelines for Data Quality Gate. Fortunately, there is a tool called Act that helps me to run my pipelines locally, without having to push code remotely. This saves me time and hassle.

The only downside of Act is that it doesn't support service containers. Link to issue

If you're working with GitLab CI, there is a similar tool called gitlab-ci-local

Did you try similar tools?
True story
Learn how to read data from a specific sheet in an Excel file using Great Expectations! My tip for today shows you how to identify the sheet name in runtime_parameters to streamline your data reading process. Check out the code snippet below to get started:
Great Expectations provides a special structure for projects, as described in the official documentation. You can create this structure through the official cli, but this requires a certain amount of user interactions. This can be problematic if you need to set up a GX project in a Docker container or AWS Lambda.

For this purpose, you can use the following code to initiate a context in a particular directory
As data quality professionals, we often need to print logs for debugging purposes. Luckily, setting up a basic configuration for logging in to Great Expectations is simple! Check out the example to get started:
If you're using SQLAlchemy and a database with Great Expectations, here's a helpful tip: You can use logging to view the queries that GX executes for validations. By enabling logging, you can gain insights into the database interactions and improve your understanding of how GX works with your data.
If you want to speed up your validations in Great Expectations, try running them in parallel. One way to do this is using #pytest, which allows you to run tests concurrently. By wrapping your checkpoints in tests and running them in parallel, you can significantly reduce the time it takes to validate your data. So give it a try and see if it improves your workflow!
#DataQuality #tips #datamanagement #GXtips #data
Why is Data Quality Important?

I have one answer which should cover all doubts about DQ's importance. Agree?
👏3
Data Management as Real Estate Management. Does it work for you?
Source
Friday !
The first course for GX. You can try GX in action without struggle with setting: Great Expectations, a data validation library for Python
True story.
Hey, do you use AWS Athena to analyze your data? If so, you might want to check out this awesome article. It shows you two ways to keep your data quality in check: AWS Glue Data Quality and Great Expectations. Looks like GX is still the best tool for Data QA
👍1
📢 Good news from Soda🎉

This week, Soda released Rule Suggestions, joining the ranks of GreatExpectation and Deequ in offering this powerful feature. 🚀 Now, you have multiple options to generate initial tests for your data source and improve your data quality management. 💪

Soda currently supports 6 rules for generating these suggestions, and you can find more information in the official documentation. 📋 Make the most out of these amazing tools and optimize your data with ease. 🌟
#DataQuality #SodaUpdates #GreatExpectation #Deequ #RuleSuggestions #Efficiency
Friday! 😜
Read the article: The new role of the Data Quality Engineer if you don't know why the Data Quality engineer is needed.
🔍 Main takeaways:

- Poor data quality costs organizations an average of $15 million per year.
- Data observability and data quality are crucial for organizations to become more data-driven and increase revenue.
- Data Quality (DQ) Engineers are needed to address data quality challenges and ensure the success of data products.
- DQ Engineers validate data flow, establish data quality metrics, and perform root cause analysis of data quality issues.
- Required skills for DQ Engineers include data management, data analysis, programming languages, data observability tools, and communication skills.
- DQ Engineers collaborate with stakeholders such as data owners, business stakeholders, data analysts, data engineering teams, and data governance teams.
- The role of DQ Engineers is to maintain high-quality data and break the cycle of "Garbage In, Garbage Out" in organizations.
🌟 Data Quality Engineers: The Key to High-Quality Data and Business Success