Data Quality
166 subscribers
63 photos
3 videos
4 files
67 links
DQ about data qa
Download Telegram
Our team recently had an in-depth brainstorming session on high-level data quality checks and SLAs, which resulted in the visual framework you see here. 🧩

This framework outlines an approach to:

- Quality Checks ✅: Applying targeted checks on raw, staged, and production data, with different priority levels for each type of check at each stage.
- Data SLAs ⏱️: Ensuring data timeliness, completeness, and adherence to schema requirements at every step.
I believe this can help you understand the initial effort required for data quality activities or improve existing processes.

Excalidraw

@data_qa - notes about data quality
🔥6❤1👍1
This media is not supported in your browser
VIEW IN TELEGRAM
As a QA specialist, I’m always glad to catch bugs, especially in data products. Today, I found a critical issue in OpenMetadata: users are not unable to disclose relationships in data lineage.

For more details, check the issue ticket here: ticket

Interestingly, this is a regression issue because older versions of OpenMetadata didn’t have this bug.

Stay tuned for updates!
👍6👀2🗿2
Hello! I mentioned the "failed rows" pattern in my previous post here. With the release of Great Expectations 1.0, there's now an expectation specifically for this pattern. UnexpectedRowsExpectation

You can enhance your tests using the power of SQL!
class ExpectConsistency(gx.expectations.UnexpectedRowsExpectation):
unexpected_rows_query: str = (
"""
select o.order_id from {batch} o
left join order_details od on o.order_id = od.order_id
where od.order_id is null
"""
)
description: str = "All records in orders should have corresponding records in order_details"

expectation = ExpectConsistency()
full_table_batch.validate(expectation)

Read official doc about UnexpectedRowsExpectation

@data_qa - notes about data quality
👍4✍1
A bit sad #Friday from @data_qa
💯5
👍4🗿4😁2
Agree ?
💯5👍3
In the world of data quality, abstraction levels continue to evolve. The journey began with writing SQL queries, moved to configuration-based tools like Soda, and now we are entering a new era of natural language interfaces.

The image illustrates this progression:

SQL: raw, powerful, but requires manual effort and experience.
Soda CL: a more user-friendly and declarative approach that simplifies data quality checks.
GenAI: the ultimate abstraction, where natural language becomes an interface. You simply describe what you need, and the system generates the logic for you.

This approach is not a silver bullet, however — it is simply a new way to interact with data. While it lowers the barrier to defining checks, it is still important to validate the results and ensure that they meet your specific needs.

Do you use ?
💯5🤔1
image_2024-12-03_20-46-57.png
290.3 KB
Image from previous post in high resolution
Data Quality at Work!

Data Stewards set the rules of the game 📜⚖️
Data Quality Engineers build the playing field 🏟🛠
Data Quality Analysts track performance on the scoreboard 📈📑👀
👍3👀2
Deprecating Columns Without Losing Data

Deprecating columns can feel more complicated than dropping entire tables, especially since they’re often part of production processes. Here’s a simple, high-level plan you can adjust to fit your needs:

Process Steps:

Mark as Deprecated: Clearly label which columns you no longer need.
Check Usage: Make sure these columns aren’t used in any jobs, reports, dashboards or ETL pipelines.
Notify Owners: Let all stakeholders know that these columns are being “deprecated.”
Plan Removal: Set up a controlled timeline to remove or isolate the column, ensuring no disruptions.

Technical Options:

Revoke Access Permissions: Remove read privileges so that queries and ETL jobs can’t touch the column.
Rename & Quarantine: Give the column a clear “off-limits” name and block access to it. This makes it easy to bring back if necessary.
Data Masking or Security Policies: Use a policy to show no real data if someone queries the column, keeping the info safe but invisible.

Key Takeaways:

1. You can “deprecate” a column without actually deleting it.
2. Restricting access or using security measures keeps data safe and recoverable.
3. A careful, step-by-step approach keeps everything running smoothly while you deprecate unnecessary columns.

Adopt these ideas, keep control, and ensure a smooth transition—even when it’s just one column at a time.
🔥4
This media is not supported in your browser
VIEW IN TELEGRAM
#Friday, and on your way home, you realize you forgot to do a crucial check in the data pipeline fix.
😁2💯2🎃1
PII vs PD

Today I faced discussion PII (Personally Identifiable Information) vs Personal data. It is not crucial topic for Data Quality but we should be familiar with difference. I think that this picture helps
👍6
🛠 Soda Checks and YAML Files: Avoid Common Pitfalls 🚀
Working with Soda checks? Here's a quick guide to avoid errors when dealing with YAML files:

Why the YAML Format Can Be Tricky
Soda checks rely heavily on YAML files, but YAML's simplicity can be deceptive:

- Even minor indentation mistakes or syntax errors can break execution.
- Pushing an invalid file can waste time and resources.
How to Avoid Issues
1. Validate YAML Syntax
Use Python’s yaml.safe_load() to verify the file is valid YAML.
Example:

import yaml
with open("your_file.yaml", "r") as file:
try:
yaml.safe_load(file)
print("YAML format is valid.")
except yaml.YAMLError as e:
print(f"YAML syntax error: {e}")

2. Check SODA-Specific Validity

Use Soda’s internal parsing method to ensure the file meets Soda’s requirements:

from soda.scan import Scan

scan = Scan()
with open("your_file.yaml", "r") as file:
yaml_str = file.read()
scan._parse_sodacl_yaml_str(yaml_str, "your_file.yaml")

if scan.has_error_logs():
print("SODA-specific errors detected!")
else:
print("SODA check file is valid.")

Full Script
For a complete solution, check out this script: Full Gist Here

By incorporating these checks, you can confidently push your YAML files without worrying about syntax or formatting errors.

Happy testing! 🎉
🔥4❤3
#Friday so true.
👍5😁3
#Friday. Happy New Year!
🥰6🎄3
As usual new year brings a lot of task but there is time for meme on #Friday
🤣8👏1💯1
#DataQuality News 📰
In a moment when the tech world is buzzing about massive investments in OpenAI and the future of AI-driven analytics, it’s more important than ever to ensure that data quality is rock-solid. After all, sophisticated AI models are only as reliable as the data they’re trained on.

Introducing DQX by Databricks Labs
If you’re working with massive datasets on Spark and Databricks, you know how crucial data quality checks are for building trust in your analytics and AI pipelines. The newly released DQX (Data Quality eX) from Databricks Labs is the latest open-source framework designed to simplify and automate data quality in Spark environments—especially for Databricks and Delta Lake users.
Key Highlights of DQX
• Databricks-Centric: Built to integrate closely with Delta Lake, Unity Catalog, and potentially Delta Live Tables, giving you a more seamless data-quality workflow on the Databricks platform.
• Scalable Checks: Leverages Spark for distributed validations, meaning it can handle large-scale datasets without slowing down your pipelines.
• Modular Design: Early documentation shows it supports easy creation of custom or out-of-the-box checks, letting you tailor validations to your data.
• Lab Status: Since it’s a Databricks Labs project, expect frequent updates and possible changes in APIs or features as it matures.

But Wait—There Are Other Options 🤚
When it comes to data quality on Spark, DQX isn’t your only choice:
Amazon Deequ ☝️
• Maturity: A widely adopted, production-ready Spark library from AWS.
• Key Feature: Boasts automated “Constraint Suggestion,” which inspects your data and suggests potential quality rules.
• Ideal For: AWS-centric teams or anyone seeking a well-documented, proven Spark-based solution.
Spark-Expectations (Nike) ❇️
• Approach: Uses a decorator pattern to validate data “in-flight” (as the Spark job runs) and “at-rest” (on existing tables).
• Key Features:
• Performs both row-level and aggregated checks in a single framework.
• Automatically quarantines records that fail one or more rules into an _error table, complete with rich metadata (failed rule details, job info, etc.).
• Provides aggregated metrics in a _stats table to avoid recalculation.
• Offers flexible action_if_failed options for each rule (e.g., fail the job, drop the row, or simply log the error).
• Ideal For: Teams needing a PySpark-based framework that does real-time or batch validations and applies immediate actions (like quarantining) based on rule failures. It’s especially useful for those who want to prevent malformed data from ever reaching downstream consumers.
While DQX could become the go-to solution for Databricks users, Amazon Deequ and Spark-Expectations remain strong alternatives—especially if you need a more mature ecosystem or a broader Spark distribution outside Databricks. As always, choose the tool that best aligns with your platform, team expertise, and specific data-quality requirements.
P.S. I mentioned only spark focused tools. For sure Soda and GX has integration with spark as well
Please open Telegram to view this post
VIEW IN TELEGRAM
👍3✍1
undoubtedly. Databricks is master to generate name for products. #Friday
👍5💯2
Not meme but future on #Friday
👍5
High-Level Data Quality Assurance (DQA) Process & Tools

Ensuring high-quality data is critical for analytics, decision-making, and ML applications. Here's a high-level DQA process along with tools for each stage:

1️⃣ Data Sources

Managed by Data Engineers, data comes from:

Data Warehouse
Data Lake
Database

2️⃣ Data Testing

Led by Data QA, this phase ensures data accuracy and reliability.

🔍 Profiling (understanding data structure, distribution, and anomalies)
🛠 Tools:
✅ PandasProfiling
✅ Soda Core

🛠 Generate Test (define validation checks for completeness, accuracy, consistency)
✅ PandasProfiling for GreatExpectations
✅ Deequ (by AWS)
✅ Hands and Brain

⚙️ Test Execution (automated tests, anomaly detection)
🛠 Tools:
✅ Soda
✅ GreatExpectations
✅ Deequ

3️⃣ Data Quality Insights

This step helps Data QA, ML, and PMs monitor and visualize data quality.

📊 Data Quality Metrics (track KPIs, identify patterns)
📈 Data Quality Dashboards (visualizing test results, reports)
🛠 Tools:
✅ Metabase
✅ SuperSet
✅ Tableau

By integrating these tools, teams can automate, monitor, and improve data quality at scale. 🚀
👍4❤1