π Netflix has open-sourced their Data/ML Workflows orchestrator Maestro. Itβs been in
It is focused on scalability and developer UX. DAGs can be defined in YAML (very similar to Databricks Jobs), Python DSL, or Java DSL.
From the post:
It is focused on scalability and developer UX. DAGs can be defined in YAML (very similar to Databricks Jobs), Python DSL, or Java DSL.
From the post:
Maestro is a general-purpose, horizontally scalable workflow orchestrator designed to manage large-scale workflows such as data pipelines and machine learning model training pipelines. It oversees the entire lifecycle of a workflow, from start to finish, including retries, queuing, task distribution to compute engines, etc.. Users can package their business logic in various formats such as Docker images, notebooks, bash script, SQL, Python, and more. Unlike traditional workflow orchestrators that only support Directed Acyclic Graphs (DAGs), Maestro supports both acyclic and cyclic workflows and also includes multiple reusable patterns, including foreach loops, subworkflow, and conditional branch, etc.
Medium
Maestro: Data/ML Workflow Orchestrator at Netflix
By Jun He, Natallia Dzenisenka, Praneeth Yenugutala, Yingyi Zhang, and Anjali Norwood
π3π€1
Data Quality
Big plans from Great Expectations team. Just a second ago GX team shared that they are going to freeze the current release on 0.18.x and release 1.0 in 4 months Breaking changes. in screenshot
greatexpectations.io
Introducing GX Core 1.0
A new name and better workflows for our open source framework
π4
This media is not supported in your browser
VIEW IN TELEGRAM
when you disable not critical test for unblocking data pipeline
#Friday
#Friday
π7
image_2024-09-17_20-27-25.png
110.3 KB
In the last two months, I minimalied incoming articles about data qa and focused on books instead of this. My opinion is that books give more structure and a deeper understanding of problems and solutions.
Today I want to recommend several books about DQ:
Data Quality Fundamentals - the book covering the main ideas about data quality. The book is good if you have nothing for DQ and you are ready to make the first step. Some ideas are out of date and relate to Monte Carlo but in general, I can recommend this book.
A step-by-step guide to improving data quality. This book is from the author of service DQOps and looks like a manual on how to set up a service. However, you can get some ideas about incident management and a dashboard view for your process.
Several days ago Anomalo published a book about Anomalo Detection "Automating Data Quality Monitoring" I had a chance to be in early subscribe to this book and have read a couple of chapters. This book is interesting because describes how to build anomaly detections (surprise ) in your data.
Now I can say that all these books are on my table and I use information from them every day. See picture
Today I want to recommend several books about DQ:
Data Quality Fundamentals - the book covering the main ideas about data quality. The book is good if you have nothing for DQ and you are ready to make the first step. Some ideas are out of date and relate to Monte Carlo but in general, I can recommend this book.
A step-by-step guide to improving data quality. This book is from the author of service DQOps and looks like a manual on how to set up a service. However, you can get some ideas about incident management and a dashboard view for your process.
Several days ago Anomalo published a book about Anomalo Detection "Automating Data Quality Monitoring" I had a chance to be in early subscribe to this book and have read a couple of chapters. This book is interesting because describes how to build anomaly detections (surprise ) in your data.
Now I can say that all these books are on my table and I use information from them every day. See picture
π6β3
π¨ Reporting Test Results with Streamlit + Soda Core π¨
Reporting is absolutely essential for presenting data quality test results. While Soda Core doesn't offer built-in reporting functionality, you can easily build a simple reporting tool using Streamlit. The process is straightforward:
Soda Core saves the scan results to a JSON file.
With Streamlit, you can read this JSON file and present it in a clean, interactive format.
π Whatβs in the Report?
Summary: Overview of the scan, including the definition name, data source, and the time taken.
Checks: Detailed info on each check (name, associated table, pass/fail outcome, diagnostics, and descriptions).
Passed/Failed Checks: Organized tables showing which checks passed and which failed.
Logs: Full logs generated during the scan.
π Select Reports: You can choose specific reports to render using a selectboxβjust pick the file you want to view.
π Sharing Reports: One of the cool things about Streamlit is the ability to share your report link with others. Share the test results with your team in just a click!
π» You can explore how easy it is to implement this in our DataQualityGate repository: GitHub Link. Check out the reporting functionality and give it a try!
π§ Also you can read about Soda-Contact-Poc in my previous post: Data Contracts PoC
Make your data quality result visibility!
Reporting is absolutely essential for presenting data quality test results. While Soda Core doesn't offer built-in reporting functionality, you can easily build a simple reporting tool using Streamlit. The process is straightforward:
Soda Core saves the scan results to a JSON file.
With Streamlit, you can read this JSON file and present it in a clean, interactive format.
π Whatβs in the Report?
Summary: Overview of the scan, including the definition name, data source, and the time taken.
Checks: Detailed info on each check (name, associated table, pass/fail outcome, diagnostics, and descriptions).
Passed/Failed Checks: Organized tables showing which checks passed and which failed.
Logs: Full logs generated during the scan.
π Select Reports: You can choose specific reports to render using a selectboxβjust pick the file you want to view.
π Sharing Reports: One of the cool things about Streamlit is the ability to share your report link with others. Share the test results with your team in just a click!
π» You can explore how easy it is to implement this in our DataQualityGate repository: GitHub Link. Check out the reporting functionality and give it a try!
π§ Also you can read about Soda-Contact-Poc in my previous post: Data Contracts PoC
Make your data quality result visibility!
GitHub
GitHub - DataQualityGate/soda-contract-poc: PoC for Soda Contracts against Vertica DB
PoC for Soda Contracts against Vertica DB . Contribute to DataQualityGate/soda-contract-poc development by creating an account on GitHub.
π₯4
Alert Fatigue
When a data quality monitoring solution fails to alert on a real issue, itβs called a false negative. When a solution triggers an alert when it shouldnβt haveβfor an issue that users donβt care about or that isnβt really a problem at allβthis is called a false positive (see image for a visual).
A system with many false positives is arguably just as problematic as a system with many false negatives because it will bombard users with unhelpful alerts, leading to the undesirable condition of alert fatigue. This is when users become so tired of responding to false alarms that they begin to ignore notifications from the system or, worse, disable notifications entirely. Itβs a bit like the platform that cried wolf. Data quality monitoring systems are particularly susceptible to alert fatigue, and itβs one of the most common reasons that adoption of a monitoring system fails.
(c) Automating Data Quality Monitoring - Anomalo
I have a new diagnosis for exhausted Data QA engineers - Alert Fatigue Syndrome
β€3π―2
Data QA and Software QA Collaboration
To enhance both Data Quality and Software Testing, collaboration between DataQA and SQA teams is crucial.
Here's why:
- Prevent Data Issues Early: Catching data quality issues before they impact production helps avoid costly mistakes and bad business decisions.
- Cross-Functional Expertise: QA teams should learn how data is used, allowing them to write more effective test cases that address both user inputs and data consistency.
- Shared Tools and Strategies: DataQA and SQA can work together using shared platforms and tools to test both application functionality and data quality. By including data checks in regular QA cycles, teams can validate that the data conforms to quality standards.
DQA Team Role:
- Provide expertise on data testing best practices.
- Recommend tools to ensure efficient data quality testing.
- Monitor and improve the effectiveness of data tests.
SQA Team Role:
- Validate user inputs according to data quality rules.
- Ensure data contract compliance and collaborate with DataQA on test design.
Together, these practices strengthen the overall testing process and improve data reliability across the company.
What do you think about collaboration?
To enhance both Data Quality and Software Testing, collaboration between DataQA and SQA teams is crucial.
Here's why:
- Prevent Data Issues Early: Catching data quality issues before they impact production helps avoid costly mistakes and bad business decisions.
- Cross-Functional Expertise: QA teams should learn how data is used, allowing them to write more effective test cases that address both user inputs and data consistency.
- Shared Tools and Strategies: DataQA and SQA can work together using shared platforms and tools to test both application functionality and data quality. By including data checks in regular QA cycles, teams can validate that the data conforms to quality standards.
DQA Team Role:
- Provide expertise on data testing best practices.
- Recommend tools to ensure efficient data quality testing.
- Monitor and improve the effectiveness of data tests.
SQA Team Role:
- Validate user inputs according to data quality rules.
- Ensure data contract compliance and collaborate with DataQA on test design.
Together, these practices strengthen the overall testing process and improve data reliability across the company.
What do you think about collaboration?
π€5π€2
This media is not supported in your browser
VIEW IN TELEGRAM
When you politely remind data producers to test their data
#Friday
#Friday
π―4π€£4π2
Our team recently had an in-depth brainstorming session on high-level data quality checks and SLAs, which resulted in the visual framework you see here. π§©
This framework outlines an approach to:
- Quality Checks β : Applying targeted checks on raw, staged, and production data, with different priority levels for each type of check at each stage.
- Data SLAs β±οΈ: Ensuring data timeliness, completeness, and adherence to schema requirements at every step.
I believe this can help you understand the initial effort required for data quality activities or improve existing processes.
Excalidraw
@data_qa - notes about data quality
This framework outlines an approach to:
- Quality Checks β : Applying targeted checks on raw, staged, and production data, with different priority levels for each type of check at each stage.
- Data SLAs β±οΈ: Ensuring data timeliness, completeness, and adherence to schema requirements at every step.
I believe this can help you understand the initial effort required for data quality activities or improve existing processes.
Excalidraw
@data_qa - notes about data quality
π₯6β€1π1
This media is not supported in your browser
VIEW IN TELEGRAM
As a QA specialist, Iβm always glad to catch bugs, especially in data products. Today, I found a critical issue in OpenMetadata: users are not unable to disclose relationships in data lineage.
For more details, check the issue ticket here: ticket
Interestingly, this is a regression issue because older versions of OpenMetadata didnβt have this bug.
Stay tuned for updates!
For more details, check the issue ticket here: ticket
Interestingly, this is a regression issue because older versions of OpenMetadata didnβt have this bug.
Stay tuned for updates!
π6π2πΏ2
Hello! I mentioned the "failed rows" pattern in my previous post here. With the release of Great Expectations 1.0, there's now an expectation specifically for this pattern. UnexpectedRowsExpectation
You can enhance your tests using the power of SQL!
Read official doc about UnexpectedRowsExpectation
@data_qa - notes about data quality
You can enhance your tests using the power of SQL!
class ExpectConsistency(gx.expectations.UnexpectedRowsExpectation):
unexpected_rows_query: str = (
"""
select o.order_id from {batch} o
left join order_details od on o.order_id = od.order_id
where od.order_id is null
"""
)
description: str = "All records in orders should have corresponding records in order_details"
expectation = ExpectConsistency()
full_table_batch.validate(expectation)
Read official doc about UnexpectedRowsExpectation
@data_qa - notes about data quality
Telegram
Data Quality
Failed Rows
Do you think about best practices in data quality checks, especially when using SQL? When you run checks against any database, you should consider the information you will receive if a test fails. The information should clearly demonstrate whyβ¦
Do you think about best practices in data quality checks, especially when using SQL? When you run checks against any database, you should consider the information you will receive if a test fails. The information should clearly demonstrate whyβ¦
π4β1
Data Quality
As a QA specialist, Iβm always glad to catch bugs, especially in data products. Today, I found a critical issue in OpenMetadata: users are not unable to disclose relationships in data lineage. For more details, check the issue ticket here: ticket Interestinglyβ¦
Today OMD notified that bug was fixed in 1.5.12. As was mentioned in duplicate ticket - this bug was added in 1.5.0
Only two learnings:
1. Your releases should be well tested
2. If you use OSS product in your company you should have staging where you will tests all new releases.
Only two learnings:
1. Your releases should be well tested
2. If you use OSS product in your company you should have staging where you will tests all new releases.
GitHub
[Lineage] User is not able to disclose relations in lineage Β· Issue #18612 Β· open-metadata/OpenMetadata
Affected module UI Describe the bug To Reproduce Table has lineage . You can add lineage manually Open lineage page for table Collapse lineage relations (click on minus) There is not ability to dis...
π4π―3
In the world of data quality, abstraction levels continue to evolve. The journey began with writing SQL queries, moved to configuration-based tools like Soda, and now we are entering a new era of natural language interfaces.
The image illustrates this progression:
SQL: raw, powerful, but requires manual effort and experience.
Soda CL: a more user-friendly and declarative approach that simplifies data quality checks.
GenAI: the ultimate abstraction, where natural language becomes an interface. You simply describe what you need, and the system generates the logic for you.
This approach is not a silver bullet, however β it is simply a new way to interact with data. While it lowers the barrier to defining checks, it is still important to validate the results and ensure that they meet your specific needs.
Do you use ?
The image illustrates this progression:
SQL: raw, powerful, but requires manual effort and experience.
Soda CL: a more user-friendly and declarative approach that simplifies data quality checks.
GenAI: the ultimate abstraction, where natural language becomes an interface. You simply describe what you need, and the system generates the logic for you.
This approach is not a silver bullet, however β it is simply a new way to interact with data. While it lowers the barrier to defining checks, it is still important to validate the results and ensure that they meet your specific needs.
Do you use ?
π―5π€1