What Is data quality tool you use?
Anonymous Poll
26%
Soda
23%
Great Expectations
3%
Deequ
17%
Internal (self-written)
29%
Just run SQL time to time
20%
Another
๐ค5
Freshness Metric in Data Quality
Introduction:
Freshness is a crucial metric in data quality that measures the timeliness and currency of data. It provides insights into how up-to-date the data is and is essential for making informed decisions and ensuring data reliability.
1. What is the Freshness Metric in Data Quality?
The freshness metric in data quality refers to how recently the data has been updated or inserted into a dataset. It indicates the time elapsed since the last update or insertion of data. A high freshness score indicates that the data is more current and reliable.
2. How to Measure Freshness?
There are different approaches to measuring freshness depending on the context and requirements. Two common methods are:
a. Business Column (updated_at): In some cases, the freshness of data is best measured using a business column, such as the "updated_at" column. This column captures the timestamp of the last update made to the data. By comparing the current timestamp with the "updated_at" value, we can determine the freshness of the data.
b. Technical Column (inserted_at): Alternatively, the freshness of data can be measured using a technical column, such as the "inserted_at" column. This column captures the timestamp when the data was inserted into the dataset. By comparing the current timestamp with the "inserted_at" value, we can assess the freshness of the data.
3. Choosing the Right Column for Freshness Measurement:
It means that you will have two types of Freshness:
1. Technical - when the table was updated
2. Business - when data was updated
Not always these params will be same
Anyway, you could always use Freshness and Latency to capture all needs in your data quality
Conclusion:
Freshness is a vital metric in data quality that helps organizations ensure the reliability and timeliness of their data. Businesses can make better decisions based on up-to-date and accurate data by understanding what freshness metric is, how to measure it, and choosing the right column for measurement.
Introduction:
Freshness is a crucial metric in data quality that measures the timeliness and currency of data. It provides insights into how up-to-date the data is and is essential for making informed decisions and ensuring data reliability.
1. What is the Freshness Metric in Data Quality?
The freshness metric in data quality refers to how recently the data has been updated or inserted into a dataset. It indicates the time elapsed since the last update or insertion of data. A high freshness score indicates that the data is more current and reliable.
2. How to Measure Freshness?
There are different approaches to measuring freshness depending on the context and requirements. Two common methods are:
a. Business Column (updated_at): In some cases, the freshness of data is best measured using a business column, such as the "updated_at" column. This column captures the timestamp of the last update made to the data. By comparing the current timestamp with the "updated_at" value, we can determine the freshness of the data.
b. Technical Column (inserted_at): Alternatively, the freshness of data can be measured using a technical column, such as the "inserted_at" column. This column captures the timestamp when the data was inserted into the dataset. By comparing the current timestamp with the "inserted_at" value, we can assess the freshness of the data.
3. Choosing the Right Column for Freshness Measurement:
It means that you will have two types of Freshness:
1. Technical - when the table was updated
2. Business - when data was updated
Not always these params will be same
Anyway, you could always use Freshness and Latency to capture all needs in your data quality
Conclusion:
Freshness is a vital metric in data quality that helps organizations ensure the reliability and timeliness of their data. Businesses can make better decisions based on up-to-date and accurate data by understanding what freshness metric is, how to measure it, and choosing the right column for measurement.
๐6๐ฏ3
GX Cloud pricing
Great Expectations published pricing for GX Cloud
They presented 3 options
- Developer - Free plan but with huge limitations. For instance, you can monitor only 5 data assets
- Team - 100$/months. 10 data assets are included and you will pay 10$/months per additional asset
- Enterprise - as usual call and make an agreement
Full plan: https://greatexpectations.io/pricing
Remember that count of data asset >= count of tables. You can have more than one data asset for one table.
Will use in your projects?
Great Expectations published pricing for GX Cloud
They presented 3 options
- Developer - Free plan but with huge limitations. For instance, you can monitor only 5 data assets
- Team - 100$/months. 10 data assets are included and you will pay 10$/months per additional asset
- Enterprise - as usual call and make an agreement
Full plan: https://greatexpectations.io/pricing
Remember that count of data asset >= count of tables. You can have more than one data asset for one table.
Will use in your projects?
๐ค3๐2
Unveiling the Layers of Data Quality Maturity
There are several approaches that help define Maturity Models. One of them is Data Quality Maturity Model by David Loshin
This model has such levels
* Initial. The process used for data quality assurance is mostly ad hoc with the most effort to respond to data quality issues.
* Repeatable. There is some management in the organization and simple information-sharing activities. There are some process disciplines, mostly it is adopted from good practice and tries to imitate the practice in the same situation.
* Defined. At this level, the team that handles data quality begins to document things like data governance policies, processes to define expectations of data quality, technology components, data quality, and report of validation processes.
* Managed. DQM includes business impact analysis, defines expectations of data quality, and measures compliance with those expectations.
* Optimized. Performance measurement across the organization can be used to identify opportunities for systemically improving data quality.
And the check is carried out on the following components
* Expectation
* Dimension
* Policy
* Procedure
* Governance
* Standardization
* Technology
* Performance Management
The assessment takes place according to a checklist and based on interviews with participants.
Then all this is summarized into results. Additionally you can make all sorts of recommendations.
Article where I got information:
[1] The Maturity Model of Data Quality Management in Banking Industry: PTXYZ Core System Customer Data
[2] Data Quality Management Maturity Measurement of Government-Owned Property Transaction in BMKG
[3] And book of David : The Practitioner's Guide to Data Quality Improvement
What is your organisation's maturity level?
There are several approaches that help define Maturity Models. One of them is Data Quality Maturity Model by David Loshin
This model has such levels
* Initial. The process used for data quality assurance is mostly ad hoc with the most effort to respond to data quality issues.
* Repeatable. There is some management in the organization and simple information-sharing activities. There are some process disciplines, mostly it is adopted from good practice and tries to imitate the practice in the same situation.
* Defined. At this level, the team that handles data quality begins to document things like data governance policies, processes to define expectations of data quality, technology components, data quality, and report of validation processes.
* Managed. DQM includes business impact analysis, defines expectations of data quality, and measures compliance with those expectations.
* Optimized. Performance measurement across the organization can be used to identify opportunities for systemically improving data quality.
And the check is carried out on the following components
* Expectation
* Dimension
* Policy
* Procedure
* Governance
* Standardization
* Technology
* Performance Management
The assessment takes place according to a checklist and based on interviews with participants.
Then all this is summarized into results. Additionally you can make all sorts of recommendations.
Article where I got information:
[1] The Maturity Model of Data Quality Management in Banking Industry: PTXYZ Core System Customer Data
[2] Data Quality Management Maturity Measurement of Government-Owned Property Transaction in BMKG
[3] And book of David : The Practitioner's Guide to Data Quality Improvement
What is your organisation's maturity level?
journal.unimma.ac.id
The Maturity Model of Data Quality Management in Banking Industry: PT XYZ Core System Customer Data | Jurnal Komtika (Komputasiโฆ
PT XYZ, engaged in the financial industry, has a target to become a leading company in Southeast Asia and has been supported by more than 200 million customer data in its core system. This huge amount of data is expected to create business opportunities,โฆ
๐5
Failed Rows
Do you think about best practices in data quality checks, especially when using SQL? When you run checks against any database, you should consider the information you will receive if a test fails. The information should clearly demonstrate why the check failed and the severity of the issue.
If the test returns FALSE, you must make additional SQL queries to understand the reason for the failure. It's crucial to grasp the cause of the check's result. The best way to implement this is by using the "failed rows" pattern. If a check fails, it returns the failed rows. These rows can include IDs from a table, allowing you to identify issues or the maximum time in a freshness check.
For instance, consider a freshness check for a table:
This query returns the maximum time if the freshness exceeds 24 hours, providing complete information about the extent of your freshness issue.
By the way, the same pattern exists in SODA CL.
Do you think about best practices in data quality checks, especially when using SQL? When you run checks against any database, you should consider the information you will receive if a test fails. The information should clearly demonstrate why the check failed and the severity of the issue.
If the test returns FALSE, you must make additional SQL queries to understand the reason for the failure. It's crucial to grasp the cause of the check's result. The best way to implement this is by using the "failed rows" pattern. If a check fails, it returns the failed rows. These rows can include IDs from a table, allowing you to identify issues or the maximum time in a freshness check.
For instance, consider a freshness check for a table:
SELECT max(ss.created_date_time) AS m_time
FROM store.sales ss
HAVING max(ss.created_date_time) < NOW() - INTERVAL '24 hours';
This query returns the maximum time if the freshness exceeds 24 hours, providing complete information about the extent of your freshness issue.
By the way, the same pattern exists in SODA CL.
๐4๐ฏ3โค1๐1
Understanding Metadata Types
There are three main types of metadata:
Business Metadata: This includes information about who owns the data, as well as definitions and descriptions of data assets. Business metadata helps provide context and meaning to data from a business perspective. It includes details like the data domain, owner information, business names, and definitions.
Technical Metadata: This type of metadata provides information about the technical aspects of data, such as column names in databases, data types, and other structural information. It describes the technical characteristics and format of data assets.
Operational Metadata: This includes information like logs and other details about how data is used, processed, and maintained over time. Operational metadata can help track data lineage, usage patterns, data quality score, and other operational aspects of data management.
Understanding these types of metadata is essential for maintaining high data quality standards and ensuring that data is accurate, reliable, and meaningful.
Do you collect metadata?
There are three main types of metadata:
Business Metadata: This includes information about who owns the data, as well as definitions and descriptions of data assets. Business metadata helps provide context and meaning to data from a business perspective. It includes details like the data domain, owner information, business names, and definitions.
Technical Metadata: This type of metadata provides information about the technical aspects of data, such as column names in databases, data types, and other structural information. It describes the technical characteristics and format of data assets.
Operational Metadata: This includes information like logs and other details about how data is used, processed, and maintained over time. Operational metadata can help track data lineage, usage patterns, data quality score, and other operational aspects of data management.
Understanding these types of metadata is essential for maintaining high data quality standards and ensuring that data is accurate, reliable, and meaningful.
Do you collect metadata?
๐ฅ4๐3โค1
๐ Netflix has open-sourced their Data/ML Workflows orchestrator Maestro. Itโs been in
It is focused on scalability and developer UX. DAGs can be defined in YAML (very similar to Databricks Jobs), Python DSL, or Java DSL.
From the post:
It is focused on scalability and developer UX. DAGs can be defined in YAML (very similar to Databricks Jobs), Python DSL, or Java DSL.
From the post:
Maestro is a general-purpose, horizontally scalable workflow orchestrator designed to manage large-scale workflows such as data pipelines and machine learning model training pipelines. It oversees the entire lifecycle of a workflow, from start to finish, including retries, queuing, task distribution to compute engines, etc.. Users can package their business logic in various formats such as Docker images, notebooks, bash script, SQL, Python, and more. Unlike traditional workflow orchestrators that only support Directed Acyclic Graphs (DAGs), Maestro supports both acyclic and cyclic workflows and also includes multiple reusable patterns, including foreach loops, subworkflow, and conditional branch, etc.
Medium
Maestro: Data/ML Workflow Orchestrator at Netflix
By Jun He, Natallia Dzenisenka, Praneeth Yenugutala, Yingyi Zhang, and Anjali Norwood
๐3๐ค1
Data Quality
Big plans from Great Expectations team. Just a second ago GX team shared that they are going to freeze the current release on 0.18.x and release 1.0 in 4 months Breaking changes. in screenshot
greatexpectations.io
Introducing GX Core 1.0
A new name and better workflows for our open source framework
๐4
This media is not supported in your browser
VIEW IN TELEGRAM
when you disable not critical test for unblocking data pipeline
#Friday
#Friday
๐7
image_2024-09-17_20-27-25.png
110.3 KB
In the last two months, I minimalied incoming articles about data qa and focused on books instead of this. My opinion is that books give more structure and a deeper understanding of problems and solutions.
Today I want to recommend several books about DQ:
Data Quality Fundamentals - the book covering the main ideas about data quality. The book is good if you have nothing for DQ and you are ready to make the first step. Some ideas are out of date and relate to Monte Carlo but in general, I can recommend this book.
A step-by-step guide to improving data quality. This book is from the author of service DQOps and looks like a manual on how to set up a service. However, you can get some ideas about incident management and a dashboard view for your process.
Several days ago Anomalo published a book about Anomalo Detection "Automating Data Quality Monitoring" I had a chance to be in early subscribe to this book and have read a couple of chapters. This book is interesting because describes how to build anomaly detections (surprise ) in your data.
Now I can say that all these books are on my table and I use information from them every day. See picture
Today I want to recommend several books about DQ:
Data Quality Fundamentals - the book covering the main ideas about data quality. The book is good if you have nothing for DQ and you are ready to make the first step. Some ideas are out of date and relate to Monte Carlo but in general, I can recommend this book.
A step-by-step guide to improving data quality. This book is from the author of service DQOps and looks like a manual on how to set up a service. However, you can get some ideas about incident management and a dashboard view for your process.
Several days ago Anomalo published a book about Anomalo Detection "Automating Data Quality Monitoring" I had a chance to be in early subscribe to this book and have read a couple of chapters. This book is interesting because describes how to build anomaly detections (surprise ) in your data.
Now I can say that all these books are on my table and I use information from them every day. See picture
๐6โ3
๐จ Reporting Test Results with Streamlit + Soda Core ๐จ
Reporting is absolutely essential for presenting data quality test results. While Soda Core doesn't offer built-in reporting functionality, you can easily build a simple reporting tool using Streamlit. The process is straightforward:
Soda Core saves the scan results to a JSON file.
With Streamlit, you can read this JSON file and present it in a clean, interactive format.
๐ Whatโs in the Report?
Summary: Overview of the scan, including the definition name, data source, and the time taken.
Checks: Detailed info on each check (name, associated table, pass/fail outcome, diagnostics, and descriptions).
Passed/Failed Checks: Organized tables showing which checks passed and which failed.
Logs: Full logs generated during the scan.
๐ Select Reports: You can choose specific reports to render using a selectboxโjust pick the file you want to view.
๐ Sharing Reports: One of the cool things about Streamlit is the ability to share your report link with others. Share the test results with your team in just a click!
๐ป You can explore how easy it is to implement this in our DataQualityGate repository: GitHub Link. Check out the reporting functionality and give it a try!
๐ง Also you can read about Soda-Contact-Poc in my previous post: Data Contracts PoC
Make your data quality result visibility!
Reporting is absolutely essential for presenting data quality test results. While Soda Core doesn't offer built-in reporting functionality, you can easily build a simple reporting tool using Streamlit. The process is straightforward:
Soda Core saves the scan results to a JSON file.
With Streamlit, you can read this JSON file and present it in a clean, interactive format.
๐ Whatโs in the Report?
Summary: Overview of the scan, including the definition name, data source, and the time taken.
Checks: Detailed info on each check (name, associated table, pass/fail outcome, diagnostics, and descriptions).
Passed/Failed Checks: Organized tables showing which checks passed and which failed.
Logs: Full logs generated during the scan.
๐ Select Reports: You can choose specific reports to render using a selectboxโjust pick the file you want to view.
๐ Sharing Reports: One of the cool things about Streamlit is the ability to share your report link with others. Share the test results with your team in just a click!
๐ป You can explore how easy it is to implement this in our DataQualityGate repository: GitHub Link. Check out the reporting functionality and give it a try!
๐ง Also you can read about Soda-Contact-Poc in my previous post: Data Contracts PoC
Make your data quality result visibility!
GitHub
GitHub - DataQualityGate/soda-contract-poc: PoC for Soda Contracts against Vertica DB
PoC for Soda Contracts against Vertica DB . Contribute to DataQualityGate/soda-contract-poc development by creating an account on GitHub.
๐ฅ4
Alert Fatigue
When a data quality monitoring solution fails to alert on a real issue, itโs called a false negative. When a solution triggers an alert when it shouldnโt haveโfor an issue that users donโt care about or that isnโt really a problem at allโthis is called a false positive (see image for a visual).
A system with many false positives is arguably just as problematic as a system with many false negatives because it will bombard users with unhelpful alerts, leading to the undesirable condition of alert fatigue. This is when users become so tired of responding to false alarms that they begin to ignore notifications from the system or, worse, disable notifications entirely. Itโs a bit like the platform that cried wolf. Data quality monitoring systems are particularly susceptible to alert fatigue, and itโs one of the most common reasons that adoption of a monitoring system fails.
(c) Automating Data Quality Monitoring - Anomalo
I have a new diagnosis for exhausted Data QA engineers - Alert Fatigue Syndrome
โค3๐ฏ2
Data QA and Software QA Collaboration
To enhance both Data Quality and Software Testing, collaboration between DataQA and SQA teams is crucial.
Here's why:
- Prevent Data Issues Early: Catching data quality issues before they impact production helps avoid costly mistakes and bad business decisions.
- Cross-Functional Expertise: QA teams should learn how data is used, allowing them to write more effective test cases that address both user inputs and data consistency.
- Shared Tools and Strategies: DataQA and SQA can work together using shared platforms and tools to test both application functionality and data quality. By including data checks in regular QA cycles, teams can validate that the data conforms to quality standards.
DQA Team Role:
- Provide expertise on data testing best practices.
- Recommend tools to ensure efficient data quality testing.
- Monitor and improve the effectiveness of data tests.
SQA Team Role:
- Validate user inputs according to data quality rules.
- Ensure data contract compliance and collaborate with DataQA on test design.
Together, these practices strengthen the overall testing process and improve data reliability across the company.
What do you think about collaboration?
To enhance both Data Quality and Software Testing, collaboration between DataQA and SQA teams is crucial.
Here's why:
- Prevent Data Issues Early: Catching data quality issues before they impact production helps avoid costly mistakes and bad business decisions.
- Cross-Functional Expertise: QA teams should learn how data is used, allowing them to write more effective test cases that address both user inputs and data consistency.
- Shared Tools and Strategies: DataQA and SQA can work together using shared platforms and tools to test both application functionality and data quality. By including data checks in regular QA cycles, teams can validate that the data conforms to quality standards.
DQA Team Role:
- Provide expertise on data testing best practices.
- Recommend tools to ensure efficient data quality testing.
- Monitor and improve the effectiveness of data tests.
SQA Team Role:
- Validate user inputs according to data quality rules.
- Ensure data contract compliance and collaborate with DataQA on test design.
Together, these practices strengthen the overall testing process and improve data reliability across the company.
What do you think about collaboration?
๐ค5๐ค2
This media is not supported in your browser
VIEW IN TELEGRAM
When you politely remind data producers to test their data
#Friday
#Friday
๐ฏ4๐คฃ4๐2
Our team recently had an in-depth brainstorming session on high-level data quality checks and SLAs, which resulted in the visual framework you see here. ๐งฉ
This framework outlines an approach to:
- Quality Checks โ : Applying targeted checks on raw, staged, and production data, with different priority levels for each type of check at each stage.
- Data SLAs โฑ๏ธ: Ensuring data timeliness, completeness, and adherence to schema requirements at every step.
I believe this can help you understand the initial effort required for data quality activities or improve existing processes.
Excalidraw
@data_qa - notes about data quality
This framework outlines an approach to:
- Quality Checks โ : Applying targeted checks on raw, staged, and production data, with different priority levels for each type of check at each stage.
- Data SLAs โฑ๏ธ: Ensuring data timeliness, completeness, and adherence to schema requirements at every step.
I believe this can help you understand the initial effort required for data quality activities or improve existing processes.
Excalidraw
@data_qa - notes about data quality
๐ฅ6โค1๐1
This media is not supported in your browser
VIEW IN TELEGRAM
As a QA specialist, Iโm always glad to catch bugs, especially in data products. Today, I found a critical issue in OpenMetadata: users are not unable to disclose relationships in data lineage.
For more details, check the issue ticket here: ticket
Interestingly, this is a regression issue because older versions of OpenMetadata didnโt have this bug.
Stay tuned for updates!
For more details, check the issue ticket here: ticket
Interestingly, this is a regression issue because older versions of OpenMetadata didnโt have this bug.
Stay tuned for updates!
๐6๐2๐ฟ2