Data Quality
166 subscribers
63 photos
3 videos
4 files
67 links
DQ about data qa
Download Telegram
Friday !
The first course for GX. You can try GX in action without struggle with setting: Great Expectations, a data validation library for Python
True story.
Hey, do you use AWS Athena to analyze your data? If so, you might want to check out this awesome article. It shows you two ways to keep your data quality in check: AWS Glue Data Quality and Great Expectations. Looks like GX is still the best tool for Data QA
๐Ÿ‘1
๐Ÿ“ข Good news from Soda๐ŸŽ‰

This week, Soda released Rule Suggestions, joining the ranks of GreatExpectation and Deequ in offering this powerful feature. ๐Ÿš€ Now, you have multiple options to generate initial tests for your data source and improve your data quality management. ๐Ÿ’ช

Soda currently supports 6 rules for generating these suggestions, and you can find more information in the official documentation. ๐Ÿ“‹ Make the most out of these amazing tools and optimize your data with ease. ๐ŸŒŸ
#DataQuality #SodaUpdates #GreatExpectation #Deequ #RuleSuggestions #Efficiency
Friday! ๐Ÿ˜œ
Read the article: The new role of the Data Quality Engineer if you don't know why the Data Quality engineer is needed.
๐Ÿ” Main takeaways:

- Poor data quality costs organizations an average of $15 million per year.
- Data observability and data quality are crucial for organizations to become more data-driven and increase revenue.
- Data Quality (DQ) Engineers are needed to address data quality challenges and ensure the success of data products.
- DQ Engineers validate data flow, establish data quality metrics, and perform root cause analysis of data quality issues.
- Required skills for DQ Engineers include data management, data analysis, programming languages, data observability tools, and communication skills.
- DQ Engineers collaborate with stakeholders such as data owners, business stakeholders, data analysts, data engineering teams, and data governance teams.
- The role of DQ Engineers is to maintain high-quality data and break the cycle of "Garbage In, Garbage Out" in organizations.
๐ŸŒŸ Data Quality Engineers: The Key to High-Quality Data and Business Success
image_2023-06-21_21-57-49.png
1.4 MB
Nice infographic about Quality Assurance in Data. My vision is to prevent bugs is cheaper and easy than bug fixes
With Great Expectations, you can easily parameterise tests, increasing coverage while maintaining a minimal test count. What's more, you have the flexibility to hide sensitive information from the test template, ensuring the privacy and security of your data.
Let's take a closer look at how it works. In your my_suite.json file, you can define your expectations using the JSON Template format. For example, you can specify the expectation to check if column values fall within a specific range look my_suite.json. Once you have defined your suite, you can run your tests with specific parameters using a validator or checkpoint.
Once you have defined your suite, you can run your tests with specific parameters using a validator or checkpoint. Look at test.py as an example.
By leveraging GX's parameterisation capabilities, you can easily customize your tests and run them with different parameter values, maximizing coverage and ensuring the integrity of your data.
So, why settle for limited test coverage or compromise data security? Take advantage of Great Expectations and unlock the full potential of your testing efforts today!
Data Quality importance in article
"One data leader I spoke with at a major transportation company told me that, on average, it took his team of 45 engineers and analysts 140 hours per week to manually check for data issues in their pipelines. Even if you have a data team of 10 people, thatโ€™s five whole days that could have spent on revenue-generating activities."
https://www.montecarlodata.com/blog-how-to-fix-your-data-quality-problem/
"According to the latest Data Quality Survey conducted by the Monte Carlo team, a significant number of respondents stated that Data Engineers are primarily responsible for Data Quality. However, I hold a different perspective on this matter; I believe that ensuring data quality is a collective responsibility that extends to the entire team. ๐Ÿ“Šโœจ #DataQuality #TeamEffort"
๐Ÿ” Data Detectives Unite! ๐Ÿ•ต๏ธโ€โ™€๏ธ๐Ÿ•ต๏ธโ€โ™‚๏ธ

When navigating the mysterious world of data, it's essential to distinguish between bugs and features. Here's your investigative roadmap:

1๏ธโƒฃ Read Documentation: Start by delving into your dataset's documentation for clues. If it's from an external source, reach out to uncover additional context. ๐Ÿ“‘

2๏ธโƒฃ Critical Questioning: Pose the pivotal query: "Is this absence intentional or an inadvertent hiccup?" ๐Ÿค”

3๏ธโƒฃ Data Doesn't Exist: If it's by design (think childless parents' child height), it's a feature, not a bug. Leave it be. ๐Ÿš€

4๏ธโƒฃ Data Wasn't Recorded: If it should've been there but isn't, you might have a sneaky bug on your hands. Investigate further. ๐Ÿž

With this flow, you'll master the art of discerning between data quirks and data quirks that need fixing. Happy data sleuthing! ๐Ÿ”๐Ÿ“Š #DataAnalysis #BugOrFeature #DataQuality
๐Ÿ‘3
Channel name was changed to ยซData Qualityยป
True story
๐Ÿ’ฏ4
Struggling with choosing the right tests for your data? I've got you covered! ๐Ÿ›ก

Here's a recommendation for your initial data test suite, addressing critical dimensions:

1. NULL Values Test: Ensure your data is free of missing or NULL values, preventing inaccuracies in your analyses.

2. Volume Tests: Essential for data quality, these tests validate row counts in crucial tables, ensuring data completeness and accuracy.

3. Uniqueness Tests: Verify the uniqueness of key data attributes to avoid duplicates that could skew your results.

4. Integrity Tests: Assess data integrity to maintain relationships between various data elements.

5. Validity Tests: Confirm data adherence to predefined rules and constraints, ensuring reliability.

6. Freshness Checks: Monitor data timeliness for informed decisions based on up-to-date information.

For further insights into these vital data quality dimensions and additional tests, refer to article: Data Quality Dimensions: Assuring Your Data Quality with Great Expectations

Don't let data quality issues hold you back. Begin with these foundational tests to ensure your data's reliability and accuracy.
#DataProfiling #DataManagement #DataAnalysis
๐Ÿ”ฅ2
Itโ€™s really crucial question.
๐Ÿ˜7โคโ€๐Ÿ”ฅ1
๐Ÿš€ Discover a Game-Changer! ๐Ÿš€

๐Ÿ”ฎ These two lines can save your life time ๐Ÿ’ฅ

๐Ÿ“Š Easily transform Ydata profiling results into a robust GreatExpectations suite with just a few lines:


profile_result = ProfileReport(df=df_ms, title="MS_Report")
profile_result.to_expectation_suite(data_context=context, run_validation=False)

๐ŸŽฏ It's that simple! Dive into more advanced cases in the article:
๐Ÿ“– Read Now

๐Ÿ›  Curious about our Data Quality Tool? It follows the same winning approach for generating technical tests. Explore repository:
๐Ÿ”— DataQualityGate

Join us on this data-driven journey to excellence! ๐ŸŒŸ #DataProfiling #DataQuality #GreatExpectations
๐Ÿ‘4โค2
๐Ÿ”๐Ÿ“Š Exploratory Data Analysis (EDA) vs. Profiling Report ๐Ÿ“ˆ๐Ÿ“‹

Let's delve into the key differences between EDA and Profiling Reports when it comes to data analysis:

Depth of Analysis:
๐Ÿ”ต EDA Notebook: EDA Notebooks are your go-to for deep analysis. They enable you to explore data with advanced statistical tests, delve into machine learning models, and conduct hypothesis testing. Analysts can dive into the data's depths to tackle complex questions head-on.

๐ŸŸ  Profiling Report: On the other hand, Profiling Reports offer a broader but less detailed view. They don't go as deep as EDA Notebooks. Instead, they swiftly summarize basic statistics and dataset characteristics.

Purpose:
๐Ÿ”ต EDA Notebook: EDA Notebooks are tailored for data analysts and scientists seeking a profound understanding. They're perfect for those wanting to thoroughly investigate data, generate hypotheses, and perform detailed exploratory analyses. EDA is the starting point for in-depth data exploration.

๐ŸŸ  Profiling Report: Profiling reports serve a different purpose. They're designed for quick data profiling, quality assessment, or sharing essential information about a dataset. These reports are handy for stakeholders who may not be data experts but need insights.

Tools to Get Started:

๐Ÿ““ Simple EDA Notebook Example: Here's a simple EDA notebook to get you started with exploratory data analysis on Kaggle.
๐Ÿ›  Tool for Data Profiling: If you're interested in data profiling, check out the ydata-profiling tool on GitHub.
Remember, the choice between EDA and Profiling Reports depends on your data analysis goals and audience! ๐Ÿ“Š๐Ÿ’ก

#DataAnalysis #ExploratoryDataAnalysis #ProfilingReport #DataExploration
๐Ÿ”ฅ3๐Ÿค”1
๐Ÿ” Data Quality Check Challenge! ๐Ÿ”

Hey there, Data Quality Enthusiasts! ๐Ÿ“Š๐Ÿ’ก

Do you trust your Data Quality Checks? ๐Ÿค” The best way to put them to the test is by introducing deliberate corruptions into your data. ๐Ÿš€ If your QA process fails to catch these issues, it's time for some improvements!

Join us for a "Test for Tests" challenge in Data Quality Assurance! ๐Ÿ“ˆ๐Ÿ” Let's ensure our QA processes are up to the mark.

#DataQuality #DataQA
โค2๐Ÿ’ฏ1
GX has an expectation called expect_table_columns_to_match_set that checks if the columns in a data frame match an unordered set. This expectation has a parameter called exact_match, which allows users to specify whether the list of columns must exactly match the observed columns. However, based on my experience, this parameter might not be clear to all users.

To provide a better understanding of how this parameter affects the expectation's behavior, let's consider an example. Assume we have a data frame called df with the following columns: "a", "b", and "c". We can implement several expectations with different values by using the GX:

pd_gx = gx.from_pandas(df)
1. pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=False)
This expectation will return True because exact_match is set to False, meaning that the data frame must have at least the expected columns (in any order).
2. pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=True)
This expectation will return False because exact_match is set to True, meaning that the data frame must have the expected columns exactly (in any order).
3. pd_gx.expect_table_columns_to_match_set(["a", "b", "c", "d"], exact_match=False)
4. pd_gx.expect_table_columns_to_match_set(["a", "b", "c", "d"], exact_match=True)
Both of these expectations will return False because the data frame does not have all the expected columns.

I believe these examples will help you gain a deeper understanding of the exact_match parameter.
๐Ÿ”ฅ4
You can try example from previos post with code snipped
import pandas as pd
import great_expectations as gx

df = pd.DataFrame({"a": [1,2,3], "b": ["a", "b", "c"], "c": [True, False, True]})
pd_gx = gx.from_pandas(df)
result = [
pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=False).success,
pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=True).success,
pd_gx.expect_table_columns_to_match_set(["a", "b","c", "d"], exact_match=False).success,
pd_gx.expect_table_columns_to_match_set(["a", "b","c", "d"], exact_match=True).success
]
result
๐Ÿ’ฏ2๐Ÿค”1๐Ÿ‘€1