DataQA requires vast experience and knowledge as it combines two areas complicated by themselves. That is the reason behind the small count of Junior DataQA specialists. They have to learn a lot before entering data testing. Do you agree?
https://towardsdatascience.com/why-data-quality-is-harder-than-code-quality-a7ab78c9d9e
https://towardsdatascience.com/why-data-quality-is-harder-than-code-quality-a7ab78c9d9e
This week I was conducting tests for our #dataquality solution Data Quality Gate (DQG). As DGQ uses AWS, I had to find a way to emulate AWS locally. I'd like to share some tools I have found to help you test your solutions w/o AWS.
1. LocalStack is a fully functional local AWS cloud stack that emulates most AWS services. I used S3 and tried to use Lambda. LS has two versions: community with s3, ec2, and others; Pro with advanced services like EKS, SageMaker, etc. I didn’t find any issues with S3 but when using Lambda I noticed that the community version supports only zip archive deployment. Deployment via docker image is available only in the Pro version. Despite it, I strongly recommend this tool for local development.
2. Moto is a similar tool to Local Stack but has some limitations in services. However, Moto can be used as a mock in Python code that allows using it for unit testing. Interesting fact: Local stack uses Moto for covering services
3. MinIO is multi-cloud object storage. It doesn’t mock AWS but can be used as S3 storage. I didn’t try this tool but I heard good reviews from my colleagues.
1. LocalStack is a fully functional local AWS cloud stack that emulates most AWS services. I used S3 and tried to use Lambda. LS has two versions: community with s3, ec2, and others; Pro with advanced services like EKS, SageMaker, etc. I didn’t find any issues with S3 but when using Lambda I noticed that the community version supports only zip archive deployment. Deployment via docker image is available only in the Pro version. Despite it, I strongly recommend this tool for local development.
2. Moto is a similar tool to Local Stack but has some limitations in services. However, Moto can be used as a mock in Python code that allows using it for unit testing. Interesting fact: Local stack uses Moto for covering services
3. MinIO is multi-cloud object storage. It doesn’t mock AWS but can be used as S3 storage. I didn’t try this tool but I heard good reviews from my colleagues.
In my experience, debugging pipelines with Jenkins was a real hassle. I would have to push any changes with every mistake and check remotely. I found the process incredibly frustrating. Nowadays, I work with GitHub Actions and create test pipelines for Data Quality Gate. Fortunately, there is a tool called Act that helps me to run my pipelines locally, without having to push code remotely. This saves me time and hassle.
The only downside of Act is that it doesn't support service containers. Link to issue
If you're working with GitLab CI, there is a similar tool called gitlab-ci-local
Did you try similar tools?
The only downside of Act is that it doesn't support service containers. Link to issue
If you're working with GitLab CI, there is a similar tool called gitlab-ci-local
Did you try similar tools?
Great Expectations provides a special structure for projects, as described in the official documentation. You can create this structure through the official cli, but this requires a certain amount of user interactions. This can be problematic if you need to set up a GX project in a Docker container or AWS Lambda.
For this purpose, you can use the following code to initiate a context in a particular directory
For this purpose, you can use the following code to initiate a context in a particular directory
The article shows how you can implement different data quality dimensions with Great Expectations. It is an important topic because Data QA s have no standard here.
Please share your feedback
Link
Please share your feedback
Link
KDnuggets
Data Quality Dimensions: Assuring Your Data Quality with Great Expectations
This article highlights the significance of ensuring high-quality data and presents six key dimensions for measuring it. These dimensions include Completeness, Consistency, Integrity, Timelessness, Uniqueness, and Validity.
If you want to speed up your validations in Great Expectations, try running them in parallel. One way to do this is using #pytest, which allows you to run tests concurrently. By wrapping your checkpoints in tests and running them in parallel, you can significantly reduce the time it takes to validate your data. So give it a try and see if it improves your workflow!
#DataQuality #tips #datamanagement #GXtips #data
#DataQuality #tips #datamanagement #GXtips #data
The first course for GX. You can try GX in action without struggle with setting: Great Expectations, a data validation library for Python
Hey, do you use AWS Athena to analyze your data? If so, you might want to check out this awesome article. It shows you two ways to keep your data quality in check: AWS Glue Data Quality and Great Expectations. Looks like GX is still the best tool for Data QA
👍1
📢 Good news from Soda🎉
This week, Soda released Rule Suggestions, joining the ranks of GreatExpectation and Deequ in offering this powerful feature. 🚀 Now, you have multiple options to generate initial tests for your data source and improve your data quality management. 💪
Soda currently supports 6 rules for generating these suggestions, and you can find more information in the official documentation. 📋 Make the most out of these amazing tools and optimize your data with ease. 🌟
#DataQuality #SodaUpdates #GreatExpectation #Deequ #RuleSuggestions #Efficiency
This week, Soda released Rule Suggestions, joining the ranks of GreatExpectation and Deequ in offering this powerful feature. 🚀 Now, you have multiple options to generate initial tests for your data source and improve your data quality management. 💪
Soda currently supports 6 rules for generating these suggestions, and you can find more information in the official documentation. 📋 Make the most out of these amazing tools and optimize your data with ease. 🌟
#DataQuality #SodaUpdates #GreatExpectation #Deequ #RuleSuggestions #Efficiency
Soda Documentation
Adopt check Suggestions
The Check Suggestions CLI assisstant is designed to simplify the process of auto-generating basic data quality checks in SodaCL.
Read the article: The new role of the Data Quality Engineer if you don't know why the Data Quality engineer is needed.
🔍 Main takeaways:
- Poor data quality costs organizations an average of $15 million per year.
- Data observability and data quality are crucial for organizations to become more data-driven and increase revenue.
- Data Quality (DQ) Engineers are needed to address data quality challenges and ensure the success of data products.
- DQ Engineers validate data flow, establish data quality metrics, and perform root cause analysis of data quality issues.
- Required skills for DQ Engineers include data management, data analysis, programming languages, data observability tools, and communication skills.
- DQ Engineers collaborate with stakeholders such as data owners, business stakeholders, data analysts, data engineering teams, and data governance teams.
- The role of DQ Engineers is to maintain high-quality data and break the cycle of "Garbage In, Garbage Out" in organizations.
🌟 Data Quality Engineers: The Key to High-Quality Data and Business Success
🔍 Main takeaways:
- Poor data quality costs organizations an average of $15 million per year.
- Data observability and data quality are crucial for organizations to become more data-driven and increase revenue.
- Data Quality (DQ) Engineers are needed to address data quality challenges and ensure the success of data products.
- DQ Engineers validate data flow, establish data quality metrics, and perform root cause analysis of data quality issues.
- Required skills for DQ Engineers include data management, data analysis, programming languages, data observability tools, and communication skills.
- DQ Engineers collaborate with stakeholders such as data owners, business stakeholders, data analysts, data engineering teams, and data governance teams.
- The role of DQ Engineers is to maintain high-quality data and break the cycle of "Garbage In, Garbage Out" in organizations.
🌟 Data Quality Engineers: The Key to High-Quality Data and Business Success
validio.io
The new role of the Data Quality Engineer
This article explores the role a Data Quality Engineer can play in solving data quality-related challenges. I also address why a DQ Engineer is needed, what value they bring to an organization, what their responsibilities could be, and the required skills…