The first course for GX. You can try GX in action without struggle with setting: Great Expectations, a data validation library for Python
Hey, do you use AWS Athena to analyze your data? If so, you might want to check out this awesome article. It shows you two ways to keep your data quality in check: AWS Glue Data Quality and Great Expectations. Looks like GX is still the best tool for Data QA
๐1
๐ข Good news from Soda๐
This week, Soda released Rule Suggestions, joining the ranks of GreatExpectation and Deequ in offering this powerful feature. ๐ Now, you have multiple options to generate initial tests for your data source and improve your data quality management. ๐ช
Soda currently supports 6 rules for generating these suggestions, and you can find more information in the official documentation. ๐ Make the most out of these amazing tools and optimize your data with ease. ๐
#DataQuality #SodaUpdates #GreatExpectation #Deequ #RuleSuggestions #Efficiency
This week, Soda released Rule Suggestions, joining the ranks of GreatExpectation and Deequ in offering this powerful feature. ๐ Now, you have multiple options to generate initial tests for your data source and improve your data quality management. ๐ช
Soda currently supports 6 rules for generating these suggestions, and you can find more information in the official documentation. ๐ Make the most out of these amazing tools and optimize your data with ease. ๐
#DataQuality #SodaUpdates #GreatExpectation #Deequ #RuleSuggestions #Efficiency
Soda Documentation
Adopt check Suggestions
The Check Suggestions CLI assisstant is designed to simplify the process of auto-generating basic data quality checks in SodaCL.
Read the article: The new role of the Data Quality Engineer if you don't know why the Data Quality engineer is needed.
๐ Main takeaways:
- Poor data quality costs organizations an average of $15 million per year.
- Data observability and data quality are crucial for organizations to become more data-driven and increase revenue.
- Data Quality (DQ) Engineers are needed to address data quality challenges and ensure the success of data products.
- DQ Engineers validate data flow, establish data quality metrics, and perform root cause analysis of data quality issues.
- Required skills for DQ Engineers include data management, data analysis, programming languages, data observability tools, and communication skills.
- DQ Engineers collaborate with stakeholders such as data owners, business stakeholders, data analysts, data engineering teams, and data governance teams.
- The role of DQ Engineers is to maintain high-quality data and break the cycle of "Garbage In, Garbage Out" in organizations.
๐ Data Quality Engineers: The Key to High-Quality Data and Business Success
๐ Main takeaways:
- Poor data quality costs organizations an average of $15 million per year.
- Data observability and data quality are crucial for organizations to become more data-driven and increase revenue.
- Data Quality (DQ) Engineers are needed to address data quality challenges and ensure the success of data products.
- DQ Engineers validate data flow, establish data quality metrics, and perform root cause analysis of data quality issues.
- Required skills for DQ Engineers include data management, data analysis, programming languages, data observability tools, and communication skills.
- DQ Engineers collaborate with stakeholders such as data owners, business stakeholders, data analysts, data engineering teams, and data governance teams.
- The role of DQ Engineers is to maintain high-quality data and break the cycle of "Garbage In, Garbage Out" in organizations.
๐ Data Quality Engineers: The Key to High-Quality Data and Business Success
validio.io
The new role of the Data Quality Engineer
This article explores the role a Data Quality Engineer can play in solving data quality-related challenges. I also address why a DQ Engineer is needed, what value they bring to an organization, what their responsibilities could be, and the required skillsโฆ
image_2023-06-21_21-57-49.png
1.4 MB
Nice infographic about Quality Assurance in Data. My vision is to prevent bugs is cheaper and easy than bug fixes
With Great Expectations, you can easily parameterise tests, increasing coverage while maintaining a minimal test count. What's more, you have the flexibility to hide sensitive information from the test template, ensuring the privacy and security of your data.
Let's take a closer look at how it works. In your my_suite.json file, you can define your expectations using the JSON Template format. For example, you can specify the expectation to check if column values fall within a specific range look my_suite.json. Once you have defined your suite, you can run your tests with specific parameters using a validator or checkpoint.
Once you have defined your suite, you can run your tests with specific parameters using a validator or checkpoint. Look at test.py as an example.
By leveraging GX's parameterisation capabilities, you can easily customize your tests and run them with different parameter values, maximizing coverage and ensuring the integrity of your data.
So, why settle for limited test coverage or compromise data security? Take advantage of Great Expectations and unlock the full potential of your testing efforts today!
Let's take a closer look at how it works. In your my_suite.json file, you can define your expectations using the JSON Template format. For example, you can specify the expectation to check if column values fall within a specific range look my_suite.json. Once you have defined your suite, you can run your tests with specific parameters using a validator or checkpoint.
Once you have defined your suite, you can run your tests with specific parameters using a validator or checkpoint. Look at test.py as an example.
By leveraging GX's parameterisation capabilities, you can easily customize your tests and run them with different parameter values, maximizing coverage and ensuring the integrity of your data.
So, why settle for limited test coverage or compromise data security? Take advantage of Great Expectations and unlock the full potential of your testing efforts today!
Data Quality importance in article
"One data leader I spoke with at a major transportation company told me that, on average, it took his team of 45 engineers and analysts 140 hours per week to manually check for data issues in their pipelines. Even if you have a data team of 10 people, thatโs five whole days that could have spent on revenue-generating activities."
https://www.montecarlodata.com/blog-how-to-fix-your-data-quality-problem/
"One data leader I spoke with at a major transportation company told me that, on average, it took his team of 45 engineers and analysts 140 hours per week to manually check for data issues in their pipelines. Even if you have a data team of 10 people, thatโs five whole days that could have spent on revenue-generating activities."
https://www.montecarlodata.com/blog-how-to-fix-your-data-quality-problem/
Monte Carlo Data
How To Fix Your Data Quality Problem
Introducing a better way to prevent bad data.
"According to the latest Data Quality Survey conducted by the Monte Carlo team, a significant number of respondents stated that Data Engineers are primarily responsible for Data Quality. However, I hold a different perspective on this matter; I believe that ensuring data quality is a collective responsibility that extends to the entire team. ๐โจ #DataQuality #TeamEffort"
๐ Data Detectives Unite! ๐ต๏ธโโ๏ธ๐ต๏ธโโ๏ธ
When navigating the mysterious world of data, it's essential to distinguish between bugs and features. Here's your investigative roadmap:
1๏ธโฃ Read Documentation: Start by delving into your dataset's documentation for clues. If it's from an external source, reach out to uncover additional context. ๐
2๏ธโฃ Critical Questioning: Pose the pivotal query: "Is this absence intentional or an inadvertent hiccup?" ๐ค
3๏ธโฃ Data Doesn't Exist: If it's by design (think childless parents' child height), it's a feature, not a bug. Leave it be. ๐
4๏ธโฃ Data Wasn't Recorded: If it should've been there but isn't, you might have a sneaky bug on your hands. Investigate further. ๐
With this flow, you'll master the art of discerning between data quirks and data quirks that need fixing. Happy data sleuthing! ๐๐ #DataAnalysis #BugOrFeature #DataQuality
When navigating the mysterious world of data, it's essential to distinguish between bugs and features. Here's your investigative roadmap:
1๏ธโฃ Read Documentation: Start by delving into your dataset's documentation for clues. If it's from an external source, reach out to uncover additional context. ๐
2๏ธโฃ Critical Questioning: Pose the pivotal query: "Is this absence intentional or an inadvertent hiccup?" ๐ค
3๏ธโฃ Data Doesn't Exist: If it's by design (think childless parents' child height), it's a feature, not a bug. Leave it be. ๐
4๏ธโฃ Data Wasn't Recorded: If it should've been there but isn't, you might have a sneaky bug on your hands. Investigate further. ๐
With this flow, you'll master the art of discerning between data quirks and data quirks that need fixing. Happy data sleuthing! ๐๐ #DataAnalysis #BugOrFeature #DataQuality
๐3
Struggling with choosing the right tests for your data? I've got you covered! ๐ก
Here's a recommendation for your initial data test suite, addressing critical dimensions:
1. NULL Values Test: Ensure your data is free of missing or NULL values, preventing inaccuracies in your analyses.
2. Volume Tests: Essential for data quality, these tests validate row counts in crucial tables, ensuring data completeness and accuracy.
3. Uniqueness Tests: Verify the uniqueness of key data attributes to avoid duplicates that could skew your results.
4. Integrity Tests: Assess data integrity to maintain relationships between various data elements.
5. Validity Tests: Confirm data adherence to predefined rules and constraints, ensuring reliability.
6. Freshness Checks: Monitor data timeliness for informed decisions based on up-to-date information.
For further insights into these vital data quality dimensions and additional tests, refer to article: Data Quality Dimensions: Assuring Your Data Quality with Great Expectations
Don't let data quality issues hold you back. Begin with these foundational tests to ensure your data's reliability and accuracy.
#DataProfiling #DataManagement #DataAnalysis
Here's a recommendation for your initial data test suite, addressing critical dimensions:
1. NULL Values Test: Ensure your data is free of missing or NULL values, preventing inaccuracies in your analyses.
2. Volume Tests: Essential for data quality, these tests validate row counts in crucial tables, ensuring data completeness and accuracy.
3. Uniqueness Tests: Verify the uniqueness of key data attributes to avoid duplicates that could skew your results.
4. Integrity Tests: Assess data integrity to maintain relationships between various data elements.
5. Validity Tests: Confirm data adherence to predefined rules and constraints, ensuring reliability.
6. Freshness Checks: Monitor data timeliness for informed decisions based on up-to-date information.
For further insights into these vital data quality dimensions and additional tests, refer to article: Data Quality Dimensions: Assuring Your Data Quality with Great Expectations
Don't let data quality issues hold you back. Begin with these foundational tests to ensure your data's reliability and accuracy.
#DataProfiling #DataManagement #DataAnalysis
๐ฅ2
๐ Discover a Game-Changer! ๐
๐ฎ These two lines can save yourlife time ๐ฅ
๐ Easily transform Ydata profiling results into a robust GreatExpectations suite with just a few lines:
๐ Read Now
๐ Curious about our Data Quality Tool? It follows the same winning approach for generating technical tests. Explore repository:
๐ DataQualityGate
Join us on this data-driven journey to excellence! ๐ #DataProfiling #DataQuality #GreatExpectations
๐ฎ These two lines can save your
๐ Easily transform Ydata profiling results into a robust GreatExpectations suite with just a few lines:
profile_result = ProfileReport(df=df_ms, title="MS_Report")๐ฏ It's that simple! Dive into more advanced cases in the article:
profile_result.to_expectation_suite(data_context=context, run_validation=False)
๐ Read Now
๐ Curious about our Data Quality Tool? It follows the same winning approach for generating technical tests. Explore repository:
๐ DataQualityGate
Join us on this data-driven journey to excellence! ๐ #DataProfiling #DataQuality #GreatExpectations
Medium
How to Use Ydata-Profiling with Great Expectations V3 API
Almost all machine learning tasks depend on data in one form or another. To generate high-quality data, data science teams needโฆ
๐4โค2
๐๐ Exploratory Data Analysis (EDA) vs. Profiling Report ๐๐
Let's delve into the key differences between EDA and Profiling Reports when it comes to data analysis:
Depth of Analysis:
๐ต EDA Notebook: EDA Notebooks are your go-to for deep analysis. They enable you to explore data with advanced statistical tests, delve into machine learning models, and conduct hypothesis testing. Analysts can dive into the data's depths to tackle complex questions head-on.
๐ Profiling Report: On the other hand, Profiling Reports offer a broader but less detailed view. They don't go as deep as EDA Notebooks. Instead, they swiftly summarize basic statistics and dataset characteristics.
Purpose:
๐ต EDA Notebook: EDA Notebooks are tailored for data analysts and scientists seeking a profound understanding. They're perfect for those wanting to thoroughly investigate data, generate hypotheses, and perform detailed exploratory analyses. EDA is the starting point for in-depth data exploration.
๐ Profiling Report: Profiling reports serve a different purpose. They're designed for quick data profiling, quality assessment, or sharing essential information about a dataset. These reports are handy for stakeholders who may not be data experts but need insights.
Tools to Get Started:
๐ Simple EDA Notebook Example: Here's a simple EDA notebook to get you started with exploratory data analysis on Kaggle.
๐ Tool for Data Profiling: If you're interested in data profiling, check out the ydata-profiling tool on GitHub.
Remember, the choice between EDA and Profiling Reports depends on your data analysis goals and audience! ๐๐ก
#DataAnalysis #ExploratoryDataAnalysis #ProfilingReport #DataExploration
Let's delve into the key differences between EDA and Profiling Reports when it comes to data analysis:
Depth of Analysis:
๐ต EDA Notebook: EDA Notebooks are your go-to for deep analysis. They enable you to explore data with advanced statistical tests, delve into machine learning models, and conduct hypothesis testing. Analysts can dive into the data's depths to tackle complex questions head-on.
๐ Profiling Report: On the other hand, Profiling Reports offer a broader but less detailed view. They don't go as deep as EDA Notebooks. Instead, they swiftly summarize basic statistics and dataset characteristics.
Purpose:
๐ต EDA Notebook: EDA Notebooks are tailored for data analysts and scientists seeking a profound understanding. They're perfect for those wanting to thoroughly investigate data, generate hypotheses, and perform detailed exploratory analyses. EDA is the starting point for in-depth data exploration.
๐ Profiling Report: Profiling reports serve a different purpose. They're designed for quick data profiling, quality assessment, or sharing essential information about a dataset. These reports are handy for stakeholders who may not be data experts but need insights.
Tools to Get Started:
๐ Simple EDA Notebook Example: Here's a simple EDA notebook to get you started with exploratory data analysis on Kaggle.
๐ Tool for Data Profiling: If you're interested in data profiling, check out the ydata-profiling tool on GitHub.
Remember, the choice between EDA and Profiling Reports depends on your data analysis goals and audience! ๐๐ก
#DataAnalysis #ExploratoryDataAnalysis #ProfilingReport #DataExploration
Kaggle
Quick Data Profiling - Data Quality Report - EDA
Explore and run machine learning code with Kaggle Notebooks | Using data from Acea Smart Water Analytics
๐ฅ3๐ค1
๐ Data Quality Check Challenge! ๐
Hey there, Data Quality Enthusiasts! ๐๐ก
Do you trust your Data Quality Checks? ๐ค The best way to put them to the test is by introducing deliberate corruptions into your data. ๐ If your QA process fails to catch these issues, it's time for some improvements!
Join us for a "Test for Tests" challenge in Data Quality Assurance! ๐๐ Let's ensure our QA processes are up to the mark.
#DataQuality #DataQA
Hey there, Data Quality Enthusiasts! ๐๐ก
Do you trust your Data Quality Checks? ๐ค The best way to put them to the test is by introducing deliberate corruptions into your data. ๐ If your QA process fails to catch these issues, it's time for some improvements!
Join us for a "Test for Tests" challenge in Data Quality Assurance! ๐๐ Let's ensure our QA processes are up to the mark.
#DataQuality #DataQA
โค2๐ฏ1
GX has an expectation called expect_table_columns_to_match_set that checks if the columns in a data frame match an unordered set. This expectation has a parameter called
To provide a better understanding of how this parameter affects the expectation's behavior, let's consider an example. Assume we have a data frame called
2.
3.
I believe these examples will help you gain a deeper understanding of the
exact_match, which allows users to specify whether the list of columns must exactly match the observed columns. However, based on my experience, this parameter might not be clear to all users.To provide a better understanding of how this parameter affects the expectation's behavior, let's consider an example. Assume we have a data frame called
df with the following columns: "a", "b", and "c". We can implement several expectations with different values by using the GX:pd_gx = gx.from_pandas(df)
1. pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=False)
This expectation will return True because exact_match is set to False, meaning that the data frame must have at least the expected columns (in any order).2.
pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=True)
This expectation will return False because exact_match is set to True, meaning that the data frame must have the expected columns exactly (in any order).3.
pd_gx.expect_table_columns_to_match_set(["a", "b", "c", "d"], exact_match=False)
4. pd_gx.expect_table_columns_to_match_set(["a", "b", "c", "d"], exact_match=True)
Both of these expectations will return False because the data frame does not have all the expected columns.I believe these examples will help you gain a deeper understanding of the
exact_match parameter.๐ฅ4
You can try example from previos post with code snipped
import pandas as pd
import great_expectations as gx
df = pd.DataFrame({"a": [1,2,3], "b": ["a", "b", "c"], "c": [True, False, True]})
pd_gx = gx.from_pandas(df)
result = [
pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=False).success,
pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=True).success,
pd_gx.expect_table_columns_to_match_set(["a", "b","c", "d"], exact_match=False).success,
pd_gx.expect_table_columns_to_match_set(["a", "b","c", "d"], exact_match=True).success
]
result
๐ฏ2๐ค1๐1