image_2024-12-03_20-46-57.png
290.3 KB
Image from previous post in high resolution
Data Quality at Work!
Data Stewards set the rules of the game ๐โ๏ธ
Data Quality Engineers build the playing field ๐๐
Data Quality Analysts track performance on the scoreboard ๐๐๐
Data Stewards set the rules of the game ๐โ๏ธ
Data Quality Engineers build the playing field ๐๐
Data Quality Analysts track performance on the scoreboard ๐๐๐
๐3๐2
Deprecating Columns Without Losing Data
Deprecating columns can feel more complicated than dropping entire tables, especially since theyโre often part of production processes. Hereโs a simple, high-level plan you can adjust to fit your needs:
Process Steps:
Mark as Deprecated: Clearly label which columns you no longer need.
Check Usage: Make sure these columns arenโt used in any jobs, reports, dashboards or ETL pipelines.
Notify Owners: Let all stakeholders know that these columns are being โdeprecated.โ
Plan Removal: Set up a controlled timeline to remove or isolate the column, ensuring no disruptions.
Technical Options:
Revoke Access Permissions: Remove read privileges so that queries and ETL jobs canโt touch the column.
Rename & Quarantine: Give the column a clear โoff-limitsโ name and block access to it. This makes it easy to bring back if necessary.
Data Masking or Security Policies: Use a policy to show no real data if someone queries the column, keeping the info safe but invisible.
Key Takeaways:
1. You can โdeprecateโ a column without actually deleting it.
2. Restricting access or using security measures keeps data safe and recoverable.
3. A careful, step-by-step approach keeps everything running smoothly while you deprecate unnecessary columns.
Adopt these ideas, keep control, and ensure a smooth transitionโeven when itโs just one column at a time.
Deprecating columns can feel more complicated than dropping entire tables, especially since theyโre often part of production processes. Hereโs a simple, high-level plan you can adjust to fit your needs:
Process Steps:
Mark as Deprecated: Clearly label which columns you no longer need.
Check Usage: Make sure these columns arenโt used in any jobs, reports, dashboards or ETL pipelines.
Notify Owners: Let all stakeholders know that these columns are being โdeprecated.โ
Plan Removal: Set up a controlled timeline to remove or isolate the column, ensuring no disruptions.
Technical Options:
Revoke Access Permissions: Remove read privileges so that queries and ETL jobs canโt touch the column.
Rename & Quarantine: Give the column a clear โoff-limitsโ name and block access to it. This makes it easy to bring back if necessary.
Data Masking or Security Policies: Use a policy to show no real data if someone queries the column, keeping the info safe but invisible.
Key Takeaways:
1. You can โdeprecateโ a column without actually deleting it.
2. Restricting access or using security measures keeps data safe and recoverable.
3. A careful, step-by-step approach keeps everything running smoothly while you deprecate unnecessary columns.
Adopt these ideas, keep control, and ensure a smooth transitionโeven when itโs just one column at a time.
๐ฅ4
This media is not supported in your browser
VIEW IN TELEGRAM
#Friday, and on your way home, you realize you forgot to do a crucial check in the data pipeline fix.
๐2๐ฏ2๐1
๐ Soda Checks and YAML Files: Avoid Common Pitfalls ๐
Working with Soda checks? Here's a quick guide to avoid errors when dealing with YAML files:
Why the YAML Format Can Be Tricky
Soda checks rely heavily on YAML files, but YAML's simplicity can be deceptive:
- Even minor indentation mistakes or syntax errors can break execution.
- Pushing an invalid file can waste time and resources.
How to Avoid Issues
1. Validate YAML Syntax
Use Pythonโs
Example:
2. Check SODA-Specific Validity
Use Sodaโs internal parsing method to ensure the file meets Sodaโs requirements:
Full Script
For a complete solution, check out this script: Full Gist Here
By incorporating these checks, you can confidently push your YAML files without worrying about syntax or formatting errors.
Happy testing! ๐
Working with Soda checks? Here's a quick guide to avoid errors when dealing with YAML files:
Why the YAML Format Can Be Tricky
Soda checks rely heavily on YAML files, but YAML's simplicity can be deceptive:
- Even minor indentation mistakes or syntax errors can break execution.
- Pushing an invalid file can waste time and resources.
How to Avoid Issues
1. Validate YAML Syntax
Use Pythonโs
yaml.safe_load() to verify the file is valid YAML.Example:
import yaml
with open("your_file.yaml", "r") as file:
try:
yaml.safe_load(file)
print("YAML format is valid.")
except yaml.YAMLError as e:
print(f"YAML syntax error: {e}")
2. Check SODA-Specific Validity
Use Sodaโs internal parsing method to ensure the file meets Sodaโs requirements:
from soda.scan import Scan
scan = Scan()
with open("your_file.yaml", "r") as file:
yaml_str = file.read()
scan._parse_sodacl_yaml_str(yaml_str, "your_file.yaml")
if scan.has_error_logs():
print("SODA-specific errors detected!")
else:
print("SODA check file is valid.")
Full Script
For a complete solution, check out this script: Full Gist Here
By incorporating these checks, you can confidently push your YAML files without worrying about syntax or formatting errors.
Happy testing! ๐
๐ฅ4โค3
#DataQuality News ๐ฐ
In a moment when the tech world is buzzing about massive investments in OpenAI and the future of AI-driven analytics, itโs more important than ever to ensure that data quality is rock-solid. After all, sophisticated AI models are only as reliable as the data theyโre trained on.
Introducing DQX by Databricks Labs
If youโre working with massive datasets on Spark and Databricks, you know how crucial data quality checks are for building trust in your analytics and AI pipelines. The newly released DQX (Data Quality eX) from Databricks Labs is the latest open-source framework designed to simplify and automate data quality in Spark environmentsโespecially for Databricks and Delta Lake users.
Key Highlights of DQX
โข Databricks-Centric: Built to integrate closely with Delta Lake, Unity Catalog, and potentially Delta Live Tables, giving you a more seamless data-quality workflow on the Databricks platform.
โข Scalable Checks: Leverages Spark for distributed validations, meaning it can handle large-scale datasets without slowing down your pipelines.
โข Modular Design: Early documentation shows it supports easy creation of custom or out-of-the-box checks, letting you tailor validations to your data.
โข Lab Status: Since itโs a Databricks Labs project, expect frequent updates and possible changes in APIs or features as it matures.
But WaitโThere Are Other Options๐ค
When it comes to data quality on Spark, DQX isnโt your only choice:
Amazon Deequโ๏ธ
โข Maturity: A widely adopted, production-ready Spark library from AWS.
โข Key Feature: Boasts automated โConstraint Suggestion,โ which inspects your data and suggests potential quality rules.
โข Ideal For: AWS-centric teams or anyone seeking a well-documented, proven Spark-based solution.
Spark-Expectations (Nike) โ๏ธ
โข Approach: Uses a decorator pattern to validate data โin-flightโ (as the Spark job runs) and โat-restโ (on existing tables).
โข Key Features:
โข Performs both row-level and aggregated checks in a single framework.
โข Automatically quarantines records that fail one or more rules into an _error table, complete with rich metadata (failed rule details, job info, etc.).
โข Provides aggregated metrics in a _stats table to avoid recalculation.
โข Offers flexible action_if_failed options for each rule (e.g., fail the job, drop the row, or simply log the error).
โข Ideal For: Teams needing a PySpark-based framework that does real-time or batch validations and applies immediate actions (like quarantining) based on rule failures. Itโs especially useful for those who want to prevent malformed data from ever reaching downstream consumers.
While DQX could become the go-to solution for Databricks users, Amazon Deequ and Spark-Expectations remain strong alternativesโespecially if you need a more mature ecosystem or a broader Spark distribution outside Databricks. As always, choose the tool that best aligns with your platform, team expertise, and specific data-quality requirements.
P.S. I mentioned only spark focused tools. For sure Soda and GX has integration with spark as well
In a moment when the tech world is buzzing about massive investments in OpenAI and the future of AI-driven analytics, itโs more important than ever to ensure that data quality is rock-solid. After all, sophisticated AI models are only as reliable as the data theyโre trained on.
Introducing DQX by Databricks Labs
If youโre working with massive datasets on Spark and Databricks, you know how crucial data quality checks are for building trust in your analytics and AI pipelines. The newly released DQX (Data Quality eX) from Databricks Labs is the latest open-source framework designed to simplify and automate data quality in Spark environmentsโespecially for Databricks and Delta Lake users.
Key Highlights of DQX
โข Databricks-Centric: Built to integrate closely with Delta Lake, Unity Catalog, and potentially Delta Live Tables, giving you a more seamless data-quality workflow on the Databricks platform.
โข Scalable Checks: Leverages Spark for distributed validations, meaning it can handle large-scale datasets without slowing down your pipelines.
โข Modular Design: Early documentation shows it supports easy creation of custom or out-of-the-box checks, letting you tailor validations to your data.
โข Lab Status: Since itโs a Databricks Labs project, expect frequent updates and possible changes in APIs or features as it matures.
But WaitโThere Are Other Options
When it comes to data quality on Spark, DQX isnโt your only choice:
Amazon Deequ
โข Maturity: A widely adopted, production-ready Spark library from AWS.
โข Key Feature: Boasts automated โConstraint Suggestion,โ which inspects your data and suggests potential quality rules.
โข Ideal For: AWS-centric teams or anyone seeking a well-documented, proven Spark-based solution.
Spark-Expectations (Nike) โ๏ธ
โข Approach: Uses a decorator pattern to validate data โin-flightโ (as the Spark job runs) and โat-restโ (on existing tables).
โข Key Features:
โข Performs both row-level and aggregated checks in a single framework.
โข Automatically quarantines records that fail one or more rules into an _error table, complete with rich metadata (failed rule details, job info, etc.).
โข Provides aggregated metrics in a _stats table to avoid recalculation.
โข Offers flexible action_if_failed options for each rule (e.g., fail the job, drop the row, or simply log the error).
โข Ideal For: Teams needing a PySpark-based framework that does real-time or batch validations and applies immediate actions (like quarantining) based on rule failures. Itโs especially useful for those who want to prevent malformed data from ever reaching downstream consumers.
While DQX could become the go-to solution for Databricks users, Amazon Deequ and Spark-Expectations remain strong alternativesโespecially if you need a more mature ecosystem or a broader Spark distribution outside Databricks. As always, choose the tool that best aligns with your platform, team expertise, and specific data-quality requirements.
P.S. I mentioned only spark focused tools. For sure Soda and GX has integration with spark as well
Please open Telegram to view this post
VIEW IN TELEGRAM
๐3โ1
High-Level Data Quality Assurance (DQA) Process & Tools
Ensuring high-quality data is critical for analytics, decision-making, and ML applications. Here's a high-level DQA process along with tools for each stage:
1๏ธโฃ Data Sources
Managed by Data Engineers, data comes from:
Data Warehouse
Data Lake
Database
2๏ธโฃ Data Testing
Led by Data QA, this phase ensures data accuracy and reliability.
๐ Profiling (understanding data structure, distribution, and anomalies)
๐ Tools:
โ PandasProfiling
โ Soda Core
๐ Generate Test (define validation checks for completeness, accuracy, consistency)
โ PandasProfiling for GreatExpectations
โ Deequ (by AWS)
โ Hands and Brain
โ๏ธ Test Execution (automated tests, anomaly detection)
๐ Tools:
โ Soda
โ GreatExpectations
โ Deequ
3๏ธโฃ Data Quality Insights
This step helps Data QA, ML, and PMs monitor and visualize data quality.
๐ Data Quality Metrics (track KPIs, identify patterns)
๐ Data Quality Dashboards (visualizing test results, reports)
๐ Tools:
โ Metabase
โ SuperSet
โ Tableau
By integrating these tools, teams can automate, monitor, and improve data quality at scale. ๐
Ensuring high-quality data is critical for analytics, decision-making, and ML applications. Here's a high-level DQA process along with tools for each stage:
1๏ธโฃ Data Sources
Managed by Data Engineers, data comes from:
Data Warehouse
Data Lake
Database
2๏ธโฃ Data Testing
Led by Data QA, this phase ensures data accuracy and reliability.
๐ Profiling (understanding data structure, distribution, and anomalies)
๐ Tools:
โ PandasProfiling
โ Soda Core
๐ Generate Test (define validation checks for completeness, accuracy, consistency)
โ PandasProfiling for GreatExpectations
โ Deequ (by AWS)
โ Hands and Brain
โ๏ธ Test Execution (automated tests, anomaly detection)
๐ Tools:
โ Soda
โ GreatExpectations
โ Deequ
3๏ธโฃ Data Quality Insights
This step helps Data QA, ML, and PMs monitor and visualize data quality.
๐ Data Quality Metrics (track KPIs, identify patterns)
๐ Data Quality Dashboards (visualizing test results, reports)
๐ Tools:
โ Metabase
โ SuperSet
โ Tableau
By integrating these tools, teams can automate, monitor, and improve data quality at scale. ๐
๐4โค1
Interesting read for all Data Quality Leaders โ I found this article on ROI that you should check out: ROI Trap by Matt Wood.
Why is it important? Just like with QA, calculating ROI for Data Quality (DQA) initiatives isnโt straightforward. Traditional metrics might miss the transformational benefits of data quality improvements โ think better decision-making, enhanced compliance, and improved customer trust.
Here are a few ways you can leverage this article to deepen your understanding of ROI for data quality:
โข Broaden Your Metrics:
Donโt limit your evaluation to direct cost savings. Consider indirect benefits like reduced data errors, improved analytics, and faster, more accurate business insights.
โข Embrace a Holistic Approach:
Just as the article illustrates with technology transitions, recognize that the true value of high-quality data goes beyond immediately measurable savings. Look at impacts on innovation, employee productivity, and strategic agility.
โข Start Small & Scale:
Consider pilot projects in key departments. Use these as case studies to build a broader ROI framework that incorporates both quantitative and qualitative benefits.
โข Shift Your Mindset:
Move from a restrictive โcost vs. savingsโ view to an enablement perspective. Understand that investing in data quality is about building a foundation for long-term business transformation.
Check out the article and let it inspire a fresh perspective on how to approach ROI in data quality!
Why is it important? Just like with QA, calculating ROI for Data Quality (DQA) initiatives isnโt straightforward. Traditional metrics might miss the transformational benefits of data quality improvements โ think better decision-making, enhanced compliance, and improved customer trust.
Here are a few ways you can leverage this article to deepen your understanding of ROI for data quality:
โข Broaden Your Metrics:
Donโt limit your evaluation to direct cost savings. Consider indirect benefits like reduced data errors, improved analytics, and faster, more accurate business insights.
โข Embrace a Holistic Approach:
Just as the article illustrates with technology transitions, recognize that the true value of high-quality data goes beyond immediately measurable savings. Look at impacts on innovation, employee productivity, and strategic agility.
โข Start Small & Scale:
Consider pilot projects in key departments. Use these as case studies to build a broader ROI framework that incorporates both quantitative and qualitative benefits.
โข Shift Your Mindset:
Move from a restrictive โcost vs. savingsโ view to an enablement perspective. Understand that investing in data quality is about building a foundation for long-term business transformation.
Check out the article and let it inspire a fresh perspective on how to approach ROI in data quality!
๐ฏ3โ1๐1
GenAI Agents are emerging as an important area in AI research, with many experts anticipating notable advancements in 2025. In this context, maintaining high data quality is essential to ensure these agents perform effectively and reliably.
Below are a few resources that provide a clear overview of GenAI Agents and emphasize the significance of robust data practices:
โข Google GenAI Agents 101 White Paper
This white paper introduces the fundamentals of GenAI Agents and discusses the role of data quality.
๐ Google GenAI Agents 101 White Paper
โข GenAI Kata
A GitHub resource that supports the design and architecture of GenAI Agents, with a focus on building strong data pipelines.
๐ GenAI Kata on GitHub
โข List of GenAI Agents
A curated list of GenAI Agents that serves as examples of various implementations
๐ List of GenAI Agents
โข Building Effective Agents from Anthropic
This resource outlines methodologies for creating effective agents
๐ Building Effective Agents
โข Crash Course on GenAI Agents
A video guide offering a clear explanation of GenAI Agents
๐ Crash Course on GenAI Agents
These resources are well-suited for anyone looking to explore the technical and practical aspects of GenAI Agents
Below are a few resources that provide a clear overview of GenAI Agents and emphasize the significance of robust data practices:
โข Google GenAI Agents 101 White Paper
This white paper introduces the fundamentals of GenAI Agents and discusses the role of data quality.
๐ Google GenAI Agents 101 White Paper
โข GenAI Kata
A GitHub resource that supports the design and architecture of GenAI Agents, with a focus on building strong data pipelines.
๐ GenAI Kata on GitHub
โข List of GenAI Agents
A curated list of GenAI Agents that serves as examples of various implementations
๐ List of GenAI Agents
โข Building Effective Agents from Anthropic
This resource outlines methodologies for creating effective agents
๐ Building Effective Agents
โข Crash Course on GenAI Agents
A video guide offering a clear explanation of GenAI Agents
๐ Crash Course on GenAI Agents
These resources are well-suited for anyone looking to explore the technical and practical aspects of GenAI Agents
๐ฅ3๐คฏ1
๐ DataHub 1.0 is Finally Here!
Great news for everyone passionate about Data Quality! DataHub officially released version 1.0, packed with powerful enhancements:
๐ฅ Improved Data Lineage Visualization
Clearly understand your data flow with enhanced visualization. Easily track data origins, transformations, and impacts, so you always trust your data.
๐ Enhanced Search Functionality
Find your datasets and metadata quicker than ever with improved search, filters, and intuitive navigation.
๐ ๏ธ Robust Metadata Management
Easily document, manage, and explore metadata with an upgraded, intuitive UI. Metadata is now more organized, accurate, and accessibleโboosting your team's confidence in data reliability.
๐ฏ Why it matters:
These powerful new features help you build trust, ensure data accuracy, and accelerate your journey to reliable, high-quality data.
Dive deeper into the full announcement: ๐ Read the full article
Great news for everyone passionate about Data Quality! DataHub officially released version 1.0, packed with powerful enhancements:
๐ฅ Improved Data Lineage Visualization
Clearly understand your data flow with enhanced visualization. Easily track data origins, transformations, and impacts, so you always trust your data.
๐ Enhanced Search Functionality
Find your datasets and metadata quicker than ever with improved search, filters, and intuitive navigation.
๐ ๏ธ Robust Metadata Management
Easily document, manage, and explore metadata with an upgraded, intuitive UI. Metadata is now more organized, accurate, and accessibleโboosting your team's confidence in data reliability.
๐ฏ Why it matters:
These powerful new features help you build trust, ensure data accuracy, and accelerate your journey to reliable, high-quality data.
Dive deeper into the full announcement: ๐ Read the full article
๐4๐1
๐ Unlocking High-Quality Dashboards at Scale โ Spotify's Approach ๐ง
Spotify shares how they ensure reliable, actionable dashboards for data-driven decision-making at scale. In 2023 alone, Spotify teams created over 4,900 dashboards, actively used by 6,000+ employees.
๐ Spotifyโs Dashboard Quality Framework:
- Vital Signs: Automated checks ensuring dashboards stay fresh and relevant.
- Spicy Dashboard Design Checklist: Manual best practices for impactful design, usability, insightfulness, and trustworthiness.
๐ Dashboards are labeled (Low, High, Golden) based on quality criteria, helping users quickly identify trustworthy insights.
๐ Spotify also developed the Dashboard Portal, a centralized internal platform enhancing dashboard discoverability and providing crucial context like ownership and update frequency.
๐ Read more
๐ฏ Interested in more on Data Quality scoring and certifications? Airbnb has an excellent deep-dive into their Data Quality Score initiative.
๐ Check out Airbnbโs article
Spotify shares how they ensure reliable, actionable dashboards for data-driven decision-making at scale. In 2023 alone, Spotify teams created over 4,900 dashboards, actively used by 6,000+ employees.
๐ Spotifyโs Dashboard Quality Framework:
- Vital Signs: Automated checks ensuring dashboards stay fresh and relevant.
- Spicy Dashboard Design Checklist: Manual best practices for impactful design, usability, insightfulness, and trustworthiness.
๐ Dashboards are labeled (Low, High, Golden) based on quality criteria, helping users quickly identify trustworthy insights.
๐ Spotify also developed the Dashboard Portal, a centralized internal platform enhancing dashboard discoverability and providing crucial context like ownership and update frequency.
๐ Read more
๐ฏ Interested in more on Data Quality scoring and certifications? Airbnb has an excellent deep-dive into their Data Quality Score initiative.
๐ Check out Airbnbโs article
Spotify Engineering
Unlocking Insights with High-Quality Dashboards at Scale | Spotify Engineering
๐ฅ4