🚨 Reporting Test Results with Streamlit + Soda Core 🚨
Reporting is absolutely essential for presenting data quality test results. While Soda Core doesn't offer built-in reporting functionality, you can easily build a simple reporting tool using Streamlit. The process is straightforward:
Soda Core saves the scan results to a JSON file.
With Streamlit, you can read this JSON file and present it in a clean, interactive format.
🔍 What’s in the Report?
Summary: Overview of the scan, including the definition name, data source, and the time taken.
Checks: Detailed info on each check (name, associated table, pass/fail outcome, diagnostics, and descriptions).
Passed/Failed Checks: Organized tables showing which checks passed and which failed.
Logs: Full logs generated during the scan.
📑 Select Reports: You can choose specific reports to render using a selectbox—just pick the file you want to view.
🔗 Sharing Reports: One of the cool things about Streamlit is the ability to share your report link with others. Share the test results with your team in just a click!
💻 You can explore how easy it is to implement this in our DataQualityGate repository: GitHub Link. Check out the reporting functionality and give it a try!
🧠 Also you can read about Soda-Contact-Poc in my previous post: Data Contracts PoC
Make your data quality result visibility!
Reporting is absolutely essential for presenting data quality test results. While Soda Core doesn't offer built-in reporting functionality, you can easily build a simple reporting tool using Streamlit. The process is straightforward:
Soda Core saves the scan results to a JSON file.
With Streamlit, you can read this JSON file and present it in a clean, interactive format.
🔍 What’s in the Report?
Summary: Overview of the scan, including the definition name, data source, and the time taken.
Checks: Detailed info on each check (name, associated table, pass/fail outcome, diagnostics, and descriptions).
Passed/Failed Checks: Organized tables showing which checks passed and which failed.
Logs: Full logs generated during the scan.
📑 Select Reports: You can choose specific reports to render using a selectbox—just pick the file you want to view.
🔗 Sharing Reports: One of the cool things about Streamlit is the ability to share your report link with others. Share the test results with your team in just a click!
💻 You can explore how easy it is to implement this in our DataQualityGate repository: GitHub Link. Check out the reporting functionality and give it a try!
🧠 Also you can read about Soda-Contact-Poc in my previous post: Data Contracts PoC
Make your data quality result visibility!
GitHub
GitHub - DataQualityGate/soda-contract-poc: PoC for Soda Contracts against Vertica DB
PoC for Soda Contracts against Vertica DB . Contribute to DataQualityGate/soda-contract-poc development by creating an account on GitHub.
🔥4
Alert Fatigue
When a data quality monitoring solution fails to alert on a real issue, it’s called a false negative. When a solution triggers an alert when it shouldn’t have—for an issue that users don’t care about or that isn’t really a problem at all—this is called a false positive (see image for a visual).
A system with many false positives is arguably just as problematic as a system with many false negatives because it will bombard users with unhelpful alerts, leading to the undesirable condition of alert fatigue. This is when users become so tired of responding to false alarms that they begin to ignore notifications from the system or, worse, disable notifications entirely. It’s a bit like the platform that cried wolf. Data quality monitoring systems are particularly susceptible to alert fatigue, and it’s one of the most common reasons that adoption of a monitoring system fails.
(c) Automating Data Quality Monitoring - Anomalo
I have a new diagnosis for exhausted Data QA engineers - Alert Fatigue Syndrome
❤3💯2
Data QA and Software QA Collaboration
To enhance both Data Quality and Software Testing, collaboration between DataQA and SQA teams is crucial.
Here's why:
- Prevent Data Issues Early: Catching data quality issues before they impact production helps avoid costly mistakes and bad business decisions.
- Cross-Functional Expertise: QA teams should learn how data is used, allowing them to write more effective test cases that address both user inputs and data consistency.
- Shared Tools and Strategies: DataQA and SQA can work together using shared platforms and tools to test both application functionality and data quality. By including data checks in regular QA cycles, teams can validate that the data conforms to quality standards.
DQA Team Role:
- Provide expertise on data testing best practices.
- Recommend tools to ensure efficient data quality testing.
- Monitor and improve the effectiveness of data tests.
SQA Team Role:
- Validate user inputs according to data quality rules.
- Ensure data contract compliance and collaborate with DataQA on test design.
Together, these practices strengthen the overall testing process and improve data reliability across the company.
What do you think about collaboration?
To enhance both Data Quality and Software Testing, collaboration between DataQA and SQA teams is crucial.
Here's why:
- Prevent Data Issues Early: Catching data quality issues before they impact production helps avoid costly mistakes and bad business decisions.
- Cross-Functional Expertise: QA teams should learn how data is used, allowing them to write more effective test cases that address both user inputs and data consistency.
- Shared Tools and Strategies: DataQA and SQA can work together using shared platforms and tools to test both application functionality and data quality. By including data checks in regular QA cycles, teams can validate that the data conforms to quality standards.
DQA Team Role:
- Provide expertise on data testing best practices.
- Recommend tools to ensure efficient data quality testing.
- Monitor and improve the effectiveness of data tests.
SQA Team Role:
- Validate user inputs according to data quality rules.
- Ensure data contract compliance and collaborate with DataQA on test design.
Together, these practices strengthen the overall testing process and improve data reliability across the company.
What do you think about collaboration?
🤝5🤔2
This media is not supported in your browser
VIEW IN TELEGRAM
When you politely remind data producers to test their data
#Friday
#Friday
💯4🤣4👍2
Our team recently had an in-depth brainstorming session on high-level data quality checks and SLAs, which resulted in the visual framework you see here. 🧩
This framework outlines an approach to:
- Quality Checks ✅: Applying targeted checks on raw, staged, and production data, with different priority levels for each type of check at each stage.
- Data SLAs ⏱️: Ensuring data timeliness, completeness, and adherence to schema requirements at every step.
I believe this can help you understand the initial effort required for data quality activities or improve existing processes.
Excalidraw
@data_qa - notes about data quality
This framework outlines an approach to:
- Quality Checks ✅: Applying targeted checks on raw, staged, and production data, with different priority levels for each type of check at each stage.
- Data SLAs ⏱️: Ensuring data timeliness, completeness, and adherence to schema requirements at every step.
I believe this can help you understand the initial effort required for data quality activities or improve existing processes.
Excalidraw
@data_qa - notes about data quality
🔥6❤1👍1
This media is not supported in your browser
VIEW IN TELEGRAM
As a QA specialist, I’m always glad to catch bugs, especially in data products. Today, I found a critical issue in OpenMetadata: users are not unable to disclose relationships in data lineage.
For more details, check the issue ticket here: ticket
Interestingly, this is a regression issue because older versions of OpenMetadata didn’t have this bug.
Stay tuned for updates!
For more details, check the issue ticket here: ticket
Interestingly, this is a regression issue because older versions of OpenMetadata didn’t have this bug.
Stay tuned for updates!
👍6👀2🗿2
Hello! I mentioned the "failed rows" pattern in my previous post here. With the release of Great Expectations 1.0, there's now an expectation specifically for this pattern. UnexpectedRowsExpectation
You can enhance your tests using the power of SQL!
Read official doc about UnexpectedRowsExpectation
@data_qa - notes about data quality
You can enhance your tests using the power of SQL!
class ExpectConsistency(gx.expectations.UnexpectedRowsExpectation):
unexpected_rows_query: str = (
"""
select o.order_id from {batch} o
left join order_details od on o.order_id = od.order_id
where od.order_id is null
"""
)
description: str = "All records in orders should have corresponding records in order_details"
expectation = ExpectConsistency()
full_table_batch.validate(expectation)
Read official doc about UnexpectedRowsExpectation
@data_qa - notes about data quality
Telegram
Data Quality
Failed Rows
Do you think about best practices in data quality checks, especially when using SQL? When you run checks against any database, you should consider the information you will receive if a test fails. The information should clearly demonstrate why…
Do you think about best practices in data quality checks, especially when using SQL? When you run checks against any database, you should consider the information you will receive if a test fails. The information should clearly demonstrate why…
👍4✍1
Data Quality
As a QA specialist, I’m always glad to catch bugs, especially in data products. Today, I found a critical issue in OpenMetadata: users are not unable to disclose relationships in data lineage. For more details, check the issue ticket here: ticket Interestingly…
Today OMD notified that bug was fixed in 1.5.12. As was mentioned in duplicate ticket - this bug was added in 1.5.0
Only two learnings:
1. Your releases should be well tested
2. If you use OSS product in your company you should have staging where you will tests all new releases.
Only two learnings:
1. Your releases should be well tested
2. If you use OSS product in your company you should have staging where you will tests all new releases.
GitHub
[Lineage] User is not able to disclose relations in lineage · Issue #18612 · open-metadata/OpenMetadata
Affected module UI Describe the bug To Reproduce Table has lineage . You can add lineage manually Open lineage page for table Collapse lineage relations (click on minus) There is not ability to dis...
👍4💯3
In the world of data quality, abstraction levels continue to evolve. The journey began with writing SQL queries, moved to configuration-based tools like Soda, and now we are entering a new era of natural language interfaces.
The image illustrates this progression:
SQL: raw, powerful, but requires manual effort and experience.
Soda CL: a more user-friendly and declarative approach that simplifies data quality checks.
GenAI: the ultimate abstraction, where natural language becomes an interface. You simply describe what you need, and the system generates the logic for you.
This approach is not a silver bullet, however — it is simply a new way to interact with data. While it lowers the barrier to defining checks, it is still important to validate the results and ensure that they meet your specific needs.
Do you use ?
The image illustrates this progression:
SQL: raw, powerful, but requires manual effort and experience.
Soda CL: a more user-friendly and declarative approach that simplifies data quality checks.
GenAI: the ultimate abstraction, where natural language becomes an interface. You simply describe what you need, and the system generates the logic for you.
This approach is not a silver bullet, however — it is simply a new way to interact with data. While it lowers the barrier to defining checks, it is still important to validate the results and ensure that they meet your specific needs.
Do you use ?
💯5🤔1
image_2024-12-03_20-46-57.png
290.3 KB
Image from previous post in high resolution
Data Quality at Work!
Data Stewards set the rules of the game 📜⚖️
Data Quality Engineers build the playing field 🏟🛠
Data Quality Analysts track performance on the scoreboard 📈📑👀
Data Stewards set the rules of the game 📜⚖️
Data Quality Engineers build the playing field 🏟🛠
Data Quality Analysts track performance on the scoreboard 📈📑👀
👍3👀2
Deprecating Columns Without Losing Data
Deprecating columns can feel more complicated than dropping entire tables, especially since they’re often part of production processes. Here’s a simple, high-level plan you can adjust to fit your needs:
Process Steps:
Mark as Deprecated: Clearly label which columns you no longer need.
Check Usage: Make sure these columns aren’t used in any jobs, reports, dashboards or ETL pipelines.
Notify Owners: Let all stakeholders know that these columns are being “deprecated.”
Plan Removal: Set up a controlled timeline to remove or isolate the column, ensuring no disruptions.
Technical Options:
Revoke Access Permissions: Remove read privileges so that queries and ETL jobs can’t touch the column.
Rename & Quarantine: Give the column a clear “off-limits” name and block access to it. This makes it easy to bring back if necessary.
Data Masking or Security Policies: Use a policy to show no real data if someone queries the column, keeping the info safe but invisible.
Key Takeaways:
1. You can “deprecate” a column without actually deleting it.
2. Restricting access or using security measures keeps data safe and recoverable.
3. A careful, step-by-step approach keeps everything running smoothly while you deprecate unnecessary columns.
Adopt these ideas, keep control, and ensure a smooth transition—even when it’s just one column at a time.
Deprecating columns can feel more complicated than dropping entire tables, especially since they’re often part of production processes. Here’s a simple, high-level plan you can adjust to fit your needs:
Process Steps:
Mark as Deprecated: Clearly label which columns you no longer need.
Check Usage: Make sure these columns aren’t used in any jobs, reports, dashboards or ETL pipelines.
Notify Owners: Let all stakeholders know that these columns are being “deprecated.”
Plan Removal: Set up a controlled timeline to remove or isolate the column, ensuring no disruptions.
Technical Options:
Revoke Access Permissions: Remove read privileges so that queries and ETL jobs can’t touch the column.
Rename & Quarantine: Give the column a clear “off-limits” name and block access to it. This makes it easy to bring back if necessary.
Data Masking or Security Policies: Use a policy to show no real data if someone queries the column, keeping the info safe but invisible.
Key Takeaways:
1. You can “deprecate” a column without actually deleting it.
2. Restricting access or using security measures keeps data safe and recoverable.
3. A careful, step-by-step approach keeps everything running smoothly while you deprecate unnecessary columns.
Adopt these ideas, keep control, and ensure a smooth transition—even when it’s just one column at a time.
🔥4
This media is not supported in your browser
VIEW IN TELEGRAM
#Friday, and on your way home, you realize you forgot to do a crucial check in the data pipeline fix.
😁2💯2🎃1
🛠 Soda Checks and YAML Files: Avoid Common Pitfalls 🚀
Working with Soda checks? Here's a quick guide to avoid errors when dealing with YAML files:
Why the YAML Format Can Be Tricky
Soda checks rely heavily on YAML files, but YAML's simplicity can be deceptive:
- Even minor indentation mistakes or syntax errors can break execution.
- Pushing an invalid file can waste time and resources.
How to Avoid Issues
1. Validate YAML Syntax
Use Python’s
Example:
2. Check SODA-Specific Validity
Use Soda’s internal parsing method to ensure the file meets Soda’s requirements:
Full Script
For a complete solution, check out this script: Full Gist Here
By incorporating these checks, you can confidently push your YAML files without worrying about syntax or formatting errors.
Happy testing! 🎉
Working with Soda checks? Here's a quick guide to avoid errors when dealing with YAML files:
Why the YAML Format Can Be Tricky
Soda checks rely heavily on YAML files, but YAML's simplicity can be deceptive:
- Even minor indentation mistakes or syntax errors can break execution.
- Pushing an invalid file can waste time and resources.
How to Avoid Issues
1. Validate YAML Syntax
Use Python’s
yaml.safe_load() to verify the file is valid YAML.Example:
import yaml
with open("your_file.yaml", "r") as file:
try:
yaml.safe_load(file)
print("YAML format is valid.")
except yaml.YAMLError as e:
print(f"YAML syntax error: {e}")
2. Check SODA-Specific Validity
Use Soda’s internal parsing method to ensure the file meets Soda’s requirements:
from soda.scan import Scan
scan = Scan()
with open("your_file.yaml", "r") as file:
yaml_str = file.read()
scan._parse_sodacl_yaml_str(yaml_str, "your_file.yaml")
if scan.has_error_logs():
print("SODA-specific errors detected!")
else:
print("SODA check file is valid.")
Full Script
For a complete solution, check out this script: Full Gist Here
By incorporating these checks, you can confidently push your YAML files without worrying about syntax or formatting errors.
Happy testing! 🎉
🔥4❤3