Data Quality
166 subscribers
63 photos
3 videos
4 files
67 links
DQ about data qa
Download Telegram
Friday! ๐Ÿ˜œ
Read the article: The new role of the Data Quality Engineer if you don't know why the Data Quality engineer is needed.
๐Ÿ” Main takeaways:

- Poor data quality costs organizations an average of $15 million per year.
- Data observability and data quality are crucial for organizations to become more data-driven and increase revenue.
- Data Quality (DQ) Engineers are needed to address data quality challenges and ensure the success of data products.
- DQ Engineers validate data flow, establish data quality metrics, and perform root cause analysis of data quality issues.
- Required skills for DQ Engineers include data management, data analysis, programming languages, data observability tools, and communication skills.
- DQ Engineers collaborate with stakeholders such as data owners, business stakeholders, data analysts, data engineering teams, and data governance teams.
- The role of DQ Engineers is to maintain high-quality data and break the cycle of "Garbage In, Garbage Out" in organizations.
๐ŸŒŸ Data Quality Engineers: The Key to High-Quality Data and Business Success
image_2023-06-21_21-57-49.png
1.4 MB
Nice infographic about Quality Assurance in Data. My vision is to prevent bugs is cheaper and easy than bug fixes
With Great Expectations, you can easily parameterise tests, increasing coverage while maintaining a minimal test count. What's more, you have the flexibility to hide sensitive information from the test template, ensuring the privacy and security of your data.
Let's take a closer look at how it works. In your my_suite.json file, you can define your expectations using the JSON Template format. For example, you can specify the expectation to check if column values fall within a specific range look my_suite.json. Once you have defined your suite, you can run your tests with specific parameters using a validator or checkpoint.
Once you have defined your suite, you can run your tests with specific parameters using a validator or checkpoint. Look at test.py as an example.
By leveraging GX's parameterisation capabilities, you can easily customize your tests and run them with different parameter values, maximizing coverage and ensuring the integrity of your data.
So, why settle for limited test coverage or compromise data security? Take advantage of Great Expectations and unlock the full potential of your testing efforts today!
Data Quality importance in article
"One data leader I spoke with at a major transportation company told me that, on average, it took his team of 45 engineers and analysts 140 hours per week to manually check for data issues in their pipelines. Even if you have a data team of 10 people, thatโ€™s five whole days that could have spent on revenue-generating activities."
https://www.montecarlodata.com/blog-how-to-fix-your-data-quality-problem/
"According to the latest Data Quality Survey conducted by the Monte Carlo team, a significant number of respondents stated that Data Engineers are primarily responsible for Data Quality. However, I hold a different perspective on this matter; I believe that ensuring data quality is a collective responsibility that extends to the entire team. ๐Ÿ“Šโœจ #DataQuality #TeamEffort"
๐Ÿ” Data Detectives Unite! ๐Ÿ•ต๏ธโ€โ™€๏ธ๐Ÿ•ต๏ธโ€โ™‚๏ธ

When navigating the mysterious world of data, it's essential to distinguish between bugs and features. Here's your investigative roadmap:

1๏ธโƒฃ Read Documentation: Start by delving into your dataset's documentation for clues. If it's from an external source, reach out to uncover additional context. ๐Ÿ“‘

2๏ธโƒฃ Critical Questioning: Pose the pivotal query: "Is this absence intentional or an inadvertent hiccup?" ๐Ÿค”

3๏ธโƒฃ Data Doesn't Exist: If it's by design (think childless parents' child height), it's a feature, not a bug. Leave it be. ๐Ÿš€

4๏ธโƒฃ Data Wasn't Recorded: If it should've been there but isn't, you might have a sneaky bug on your hands. Investigate further. ๐Ÿž

With this flow, you'll master the art of discerning between data quirks and data quirks that need fixing. Happy data sleuthing! ๐Ÿ”๐Ÿ“Š #DataAnalysis #BugOrFeature #DataQuality
๐Ÿ‘3
Channel name was changed to ยซData Qualityยป
True story
๐Ÿ’ฏ4
Struggling with choosing the right tests for your data? I've got you covered! ๐Ÿ›ก

Here's a recommendation for your initial data test suite, addressing critical dimensions:

1. NULL Values Test: Ensure your data is free of missing or NULL values, preventing inaccuracies in your analyses.

2. Volume Tests: Essential for data quality, these tests validate row counts in crucial tables, ensuring data completeness and accuracy.

3. Uniqueness Tests: Verify the uniqueness of key data attributes to avoid duplicates that could skew your results.

4. Integrity Tests: Assess data integrity to maintain relationships between various data elements.

5. Validity Tests: Confirm data adherence to predefined rules and constraints, ensuring reliability.

6. Freshness Checks: Monitor data timeliness for informed decisions based on up-to-date information.

For further insights into these vital data quality dimensions and additional tests, refer to article: Data Quality Dimensions: Assuring Your Data Quality with Great Expectations

Don't let data quality issues hold you back. Begin with these foundational tests to ensure your data's reliability and accuracy.
#DataProfiling #DataManagement #DataAnalysis
๐Ÿ”ฅ2
Itโ€™s really crucial question.
๐Ÿ˜7โคโ€๐Ÿ”ฅ1
๐Ÿš€ Discover a Game-Changer! ๐Ÿš€

๐Ÿ”ฎ These two lines can save your life time ๐Ÿ’ฅ

๐Ÿ“Š Easily transform Ydata profiling results into a robust GreatExpectations suite with just a few lines:


profile_result = ProfileReport(df=df_ms, title="MS_Report")
profile_result.to_expectation_suite(data_context=context, run_validation=False)

๐ŸŽฏ It's that simple! Dive into more advanced cases in the article:
๐Ÿ“– Read Now

๐Ÿ›  Curious about our Data Quality Tool? It follows the same winning approach for generating technical tests. Explore repository:
๐Ÿ”— DataQualityGate

Join us on this data-driven journey to excellence! ๐ŸŒŸ #DataProfiling #DataQuality #GreatExpectations
๐Ÿ‘4โค2
๐Ÿ”๐Ÿ“Š Exploratory Data Analysis (EDA) vs. Profiling Report ๐Ÿ“ˆ๐Ÿ“‹

Let's delve into the key differences between EDA and Profiling Reports when it comes to data analysis:

Depth of Analysis:
๐Ÿ”ต EDA Notebook: EDA Notebooks are your go-to for deep analysis. They enable you to explore data with advanced statistical tests, delve into machine learning models, and conduct hypothesis testing. Analysts can dive into the data's depths to tackle complex questions head-on.

๐ŸŸ  Profiling Report: On the other hand, Profiling Reports offer a broader but less detailed view. They don't go as deep as EDA Notebooks. Instead, they swiftly summarize basic statistics and dataset characteristics.

Purpose:
๐Ÿ”ต EDA Notebook: EDA Notebooks are tailored for data analysts and scientists seeking a profound understanding. They're perfect for those wanting to thoroughly investigate data, generate hypotheses, and perform detailed exploratory analyses. EDA is the starting point for in-depth data exploration.

๐ŸŸ  Profiling Report: Profiling reports serve a different purpose. They're designed for quick data profiling, quality assessment, or sharing essential information about a dataset. These reports are handy for stakeholders who may not be data experts but need insights.

Tools to Get Started:

๐Ÿ““ Simple EDA Notebook Example: Here's a simple EDA notebook to get you started with exploratory data analysis on Kaggle.
๐Ÿ›  Tool for Data Profiling: If you're interested in data profiling, check out the ydata-profiling tool on GitHub.
Remember, the choice between EDA and Profiling Reports depends on your data analysis goals and audience! ๐Ÿ“Š๐Ÿ’ก

#DataAnalysis #ExploratoryDataAnalysis #ProfilingReport #DataExploration
๐Ÿ”ฅ3๐Ÿค”1
๐Ÿ” Data Quality Check Challenge! ๐Ÿ”

Hey there, Data Quality Enthusiasts! ๐Ÿ“Š๐Ÿ’ก

Do you trust your Data Quality Checks? ๐Ÿค” The best way to put them to the test is by introducing deliberate corruptions into your data. ๐Ÿš€ If your QA process fails to catch these issues, it's time for some improvements!

Join us for a "Test for Tests" challenge in Data Quality Assurance! ๐Ÿ“ˆ๐Ÿ” Let's ensure our QA processes are up to the mark.

#DataQuality #DataQA
โค2๐Ÿ’ฏ1
GX has an expectation called expect_table_columns_to_match_set that checks if the columns in a data frame match an unordered set. This expectation has a parameter called exact_match, which allows users to specify whether the list of columns must exactly match the observed columns. However, based on my experience, this parameter might not be clear to all users.

To provide a better understanding of how this parameter affects the expectation's behavior, let's consider an example. Assume we have a data frame called df with the following columns: "a", "b", and "c". We can implement several expectations with different values by using the GX:

pd_gx = gx.from_pandas(df)
1. pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=False)
This expectation will return True because exact_match is set to False, meaning that the data frame must have at least the expected columns (in any order).
2. pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=True)
This expectation will return False because exact_match is set to True, meaning that the data frame must have the expected columns exactly (in any order).
3. pd_gx.expect_table_columns_to_match_set(["a", "b", "c", "d"], exact_match=False)
4. pd_gx.expect_table_columns_to_match_set(["a", "b", "c", "d"], exact_match=True)
Both of these expectations will return False because the data frame does not have all the expected columns.

I believe these examples will help you gain a deeper understanding of the exact_match parameter.
๐Ÿ”ฅ4
You can try example from previos post with code snipped
import pandas as pd
import great_expectations as gx

df = pd.DataFrame({"a": [1,2,3], "b": ["a", "b", "c"], "c": [True, False, True]})
pd_gx = gx.from_pandas(df)
result = [
pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=False).success,
pd_gx.expect_table_columns_to_match_set(["a", "b"], exact_match=True).success,
pd_gx.expect_table_columns_to_match_set(["a", "b","c", "d"], exact_match=False).success,
pd_gx.expect_table_columns_to_match_set(["a", "b","c", "d"], exact_match=True).success
]
result
๐Ÿ’ฏ2๐Ÿค”1๐Ÿ‘€1
๐Ÿคฃ4๐Ÿคฏ1
# Soda Quick Start

Get ready for data validation on Vertica using Soda! The key point: Soda has two package types:

- soda-vertica: The Enterprise package.
- soda-core-vertica: The OSS package.

Find OSS documentation on GitHub [here](https://github.com/sodadata/soda-core/blob/main/docs/configuration.md). Installing both? Disaster. Keep those packages in your environment ๐Ÿคฆโ€โ™‚๏ธ. Control everything with this command:

pip freeze | grep soda 
soda-core==3.0.53
soda-core-vertica==3.0.53


Now, craft a config file like this:

data_source vertica_local:
type: vertica
connection:
host: localhost
port: '5433'
username: dbadmin
password: foo123
database: Vmart
schema: online_sales


Create a file for checks, using built-in or SQL. Set expectations for SQL results, like expecting failed rows:

Checks for online_sales.online_page_dimension:
- failed rows:
fail query: |
SELECT * FROM online_sales.online_page_dimension
attributes:
department: Sales
priority: 1
tags: [completeness]


Now, launch the command:

soda scan -d vertica_local -c configuration.yml -srf test-local.json local-checks.yml


Results in the log and a file with a treasure trove of useful info in JSON! ๐Ÿš€
๐Ÿ‘3๐Ÿ”ฅ1
Introducing "Spark-Expectations" a versatile data quality tool, inspired by DLT, that offers real-time and at-rest data quality checks, transforming the way data quality is managed. Its standout features include:

Real-Time Data Quality: It performs data quality checks in real-time during Spark job execution, ensuring data quality from start to finish.

Holistic Data Quality: Beyond real-time checks, it also assesses data quality in at-rest data, covering the entire data lifecycle.

Tackling Common Data Quality Challenges:

Handling Malformed Data: Unlike other tools, it can automatically remove problematic data from the original dataset.

Comprehensive Checks: It uniquely combines row and column-level data quality checks in one tool, simplifying the process.

Automated Error Detection: Faulty records are isolated in an error table, simplifying collaboration to rectify issues.

Efficient Error Handling: By segregating erroneous data, it eases downstream processing and correction.

Streamlined Correction: It simplifies the process of addressing data errors, reducing planning requirements.

Key Principles of Spark-Expectations:

Detailed Error Reporting: It supports individual row-based and aggregated data quality checks, providing comprehensive error information.

Aggregated Metrics: Offers aggregated metrics at the job level, reducing the need for extensive recalculations.

Selective Data Writing: Data that fails quality checks isn't written to the final table by default, preventing error propagation.

Flexible Notifications: Users can set notifications at job start, completion, or on failure.

Customizable Error Handling: It allows users to define actions if a rule fails.

In summary, Spark-Expectations is a robust data quality tool that simplifies error handling, automates data quality checks, and enhances data quality management throughout the data lifecycle.

Did you compare Spark Expectations with Deequ?
๐Ÿ‘5๐Ÿ‘€2
Big plans from Great Expectations team.
Just a second ago GX team shared that they are going to freeze the current release on 0.18.x and release 1.0 in 4 months

Breaking changes. in screenshot
๐Ÿ‘4๐Ÿ‘€4๐Ÿ˜ฑ2
Data Contracts are a crucial part of data pipelines. The availability features for data contract validation in data quality tools represent the killer feature. One month ago, Soda released Data Contract Validation. You can use it to enforce various aspects, such as a datasetโ€™s schema, the data types of columns, unique values in a column, acceptable values in a column, completeness of a column, valid data referenced against another set of data, and more!

For easy onboarding, the Data Quality Gate team prepared a Proof of Concept (POC) to demonstrate the power of data contracts:
- Connect to Vertica DB
- Generate data contracts based on profiling
- Perform checks using Soda SQL

Try it out and return with feedback: GitHub Link
๐Ÿ”ฅ6โคโ€๐Ÿ”ฅ2