High-Level Data Quality Assurance (DQA) Process & Tools
Ensuring high-quality data is critical for analytics, decision-making, and ML applications. Here's a high-level DQA process along with tools for each stage:
1️⃣ Data Sources
Managed by Data Engineers, data comes from:
Data Warehouse
Data Lake
Database
2️⃣ Data Testing
Led by Data QA, this phase ensures data accuracy and reliability.
🔍 Profiling (understanding data structure, distribution, and anomalies)
🛠 Tools:
✅ PandasProfiling
✅ Soda Core
🛠 Generate Test (define validation checks for completeness, accuracy, consistency)
✅ PandasProfiling for GreatExpectations
✅ Deequ (by AWS)
✅ Hands and Brain
⚙️ Test Execution (automated tests, anomaly detection)
🛠 Tools:
✅ Soda
✅ GreatExpectations
✅ Deequ
3️⃣ Data Quality Insights
This step helps Data QA, ML, and PMs monitor and visualize data quality.
📊 Data Quality Metrics (track KPIs, identify patterns)
📈 Data Quality Dashboards (visualizing test results, reports)
🛠 Tools:
✅ Metabase
✅ SuperSet
✅ Tableau
By integrating these tools, teams can automate, monitor, and improve data quality at scale. 🚀
Ensuring high-quality data is critical for analytics, decision-making, and ML applications. Here's a high-level DQA process along with tools for each stage:
1️⃣ Data Sources
Managed by Data Engineers, data comes from:
Data Warehouse
Data Lake
Database
2️⃣ Data Testing
Led by Data QA, this phase ensures data accuracy and reliability.
🔍 Profiling (understanding data structure, distribution, and anomalies)
🛠 Tools:
✅ PandasProfiling
✅ Soda Core
🛠 Generate Test (define validation checks for completeness, accuracy, consistency)
✅ PandasProfiling for GreatExpectations
✅ Deequ (by AWS)
✅ Hands and Brain
⚙️ Test Execution (automated tests, anomaly detection)
🛠 Tools:
✅ Soda
✅ GreatExpectations
✅ Deequ
3️⃣ Data Quality Insights
This step helps Data QA, ML, and PMs monitor and visualize data quality.
📊 Data Quality Metrics (track KPIs, identify patterns)
📈 Data Quality Dashboards (visualizing test results, reports)
🛠 Tools:
✅ Metabase
✅ SuperSet
✅ Tableau
By integrating these tools, teams can automate, monitor, and improve data quality at scale. 🚀
👍4❤1
Interesting read for all Data Quality Leaders – I found this article on ROI that you should check out: ROI Trap by Matt Wood.
Why is it important? Just like with QA, calculating ROI for Data Quality (DQA) initiatives isn’t straightforward. Traditional metrics might miss the transformational benefits of data quality improvements – think better decision-making, enhanced compliance, and improved customer trust.
Here are a few ways you can leverage this article to deepen your understanding of ROI for data quality:
• Broaden Your Metrics:
Don’t limit your evaluation to direct cost savings. Consider indirect benefits like reduced data errors, improved analytics, and faster, more accurate business insights.
• Embrace a Holistic Approach:
Just as the article illustrates with technology transitions, recognize that the true value of high-quality data goes beyond immediately measurable savings. Look at impacts on innovation, employee productivity, and strategic agility.
• Start Small & Scale:
Consider pilot projects in key departments. Use these as case studies to build a broader ROI framework that incorporates both quantitative and qualitative benefits.
• Shift Your Mindset:
Move from a restrictive “cost vs. savings” view to an enablement perspective. Understand that investing in data quality is about building a foundation for long-term business transformation.
Check out the article and let it inspire a fresh perspective on how to approach ROI in data quality!
Why is it important? Just like with QA, calculating ROI for Data Quality (DQA) initiatives isn’t straightforward. Traditional metrics might miss the transformational benefits of data quality improvements – think better decision-making, enhanced compliance, and improved customer trust.
Here are a few ways you can leverage this article to deepen your understanding of ROI for data quality:
• Broaden Your Metrics:
Don’t limit your evaluation to direct cost savings. Consider indirect benefits like reduced data errors, improved analytics, and faster, more accurate business insights.
• Embrace a Holistic Approach:
Just as the article illustrates with technology transitions, recognize that the true value of high-quality data goes beyond immediately measurable savings. Look at impacts on innovation, employee productivity, and strategic agility.
• Start Small & Scale:
Consider pilot projects in key departments. Use these as case studies to build a broader ROI framework that incorporates both quantitative and qualitative benefits.
• Shift Your Mindset:
Move from a restrictive “cost vs. savings” view to an enablement perspective. Understand that investing in data quality is about building a foundation for long-term business transformation.
Check out the article and let it inspire a fresh perspective on how to approach ROI in data quality!
💯3✍1👍1
GenAI Agents are emerging as an important area in AI research, with many experts anticipating notable advancements in 2025. In this context, maintaining high data quality is essential to ensure these agents perform effectively and reliably.
Below are a few resources that provide a clear overview of GenAI Agents and emphasize the significance of robust data practices:
• Google GenAI Agents 101 White Paper
This white paper introduces the fundamentals of GenAI Agents and discusses the role of data quality.
🔗 Google GenAI Agents 101 White Paper
• GenAI Kata
A GitHub resource that supports the design and architecture of GenAI Agents, with a focus on building strong data pipelines.
🔗 GenAI Kata on GitHub
• List of GenAI Agents
A curated list of GenAI Agents that serves as examples of various implementations
🔗 List of GenAI Agents
• Building Effective Agents from Anthropic
This resource outlines methodologies for creating effective agents
🔗 Building Effective Agents
• Crash Course on GenAI Agents
A video guide offering a clear explanation of GenAI Agents
🔗 Crash Course on GenAI Agents
These resources are well-suited for anyone looking to explore the technical and practical aspects of GenAI Agents
Below are a few resources that provide a clear overview of GenAI Agents and emphasize the significance of robust data practices:
• Google GenAI Agents 101 White Paper
This white paper introduces the fundamentals of GenAI Agents and discusses the role of data quality.
🔗 Google GenAI Agents 101 White Paper
• GenAI Kata
A GitHub resource that supports the design and architecture of GenAI Agents, with a focus on building strong data pipelines.
🔗 GenAI Kata on GitHub
• List of GenAI Agents
A curated list of GenAI Agents that serves as examples of various implementations
🔗 List of GenAI Agents
• Building Effective Agents from Anthropic
This resource outlines methodologies for creating effective agents
🔗 Building Effective Agents
• Crash Course on GenAI Agents
A video guide offering a clear explanation of GenAI Agents
🔗 Crash Course on GenAI Agents
These resources are well-suited for anyone looking to explore the technical and practical aspects of GenAI Agents
🔥3🤯1
🚀 DataHub 1.0 is Finally Here!
Great news for everyone passionate about Data Quality! DataHub officially released version 1.0, packed with powerful enhancements:
🔥 Improved Data Lineage Visualization
Clearly understand your data flow with enhanced visualization. Easily track data origins, transformations, and impacts, so you always trust your data.
🔍 Enhanced Search Functionality
Find your datasets and metadata quicker than ever with improved search, filters, and intuitive navigation.
🛠️ Robust Metadata Management
Easily document, manage, and explore metadata with an upgraded, intuitive UI. Metadata is now more organized, accurate, and accessible—boosting your team's confidence in data reliability.
🎯 Why it matters:
These powerful new features help you build trust, ensure data accuracy, and accelerate your journey to reliable, high-quality data.
Dive deeper into the full announcement: 👉 Read the full article
Great news for everyone passionate about Data Quality! DataHub officially released version 1.0, packed with powerful enhancements:
🔥 Improved Data Lineage Visualization
Clearly understand your data flow with enhanced visualization. Easily track data origins, transformations, and impacts, so you always trust your data.
🔍 Enhanced Search Functionality
Find your datasets and metadata quicker than ever with improved search, filters, and intuitive navigation.
🛠️ Robust Metadata Management
Easily document, manage, and explore metadata with an upgraded, intuitive UI. Metadata is now more organized, accurate, and accessible—boosting your team's confidence in data reliability.
🎯 Why it matters:
These powerful new features help you build trust, ensure data accuracy, and accelerate your journey to reliable, high-quality data.
Dive deeper into the full announcement: 👉 Read the full article
👍4👏1
🚀 Unlocking High-Quality Dashboards at Scale – Spotify's Approach 🎧
Spotify shares how they ensure reliable, actionable dashboards for data-driven decision-making at scale. In 2023 alone, Spotify teams created over 4,900 dashboards, actively used by 6,000+ employees.
🔑 Spotify’s Dashboard Quality Framework:
- Vital Signs: Automated checks ensuring dashboards stay fresh and relevant.
- Spicy Dashboard Design Checklist: Manual best practices for impactful design, usability, insightfulness, and trustworthiness.
🏅 Dashboards are labeled (Low, High, Golden) based on quality criteria, helping users quickly identify trustworthy insights.
📌 Spotify also developed the Dashboard Portal, a centralized internal platform enhancing dashboard discoverability and providing crucial context like ownership and update frequency.
👉 Read more
🎯 Interested in more on Data Quality scoring and certifications? Airbnb has an excellent deep-dive into their Data Quality Score initiative.
👉 Check out Airbnb’s article
Spotify shares how they ensure reliable, actionable dashboards for data-driven decision-making at scale. In 2023 alone, Spotify teams created over 4,900 dashboards, actively used by 6,000+ employees.
🔑 Spotify’s Dashboard Quality Framework:
- Vital Signs: Automated checks ensuring dashboards stay fresh and relevant.
- Spicy Dashboard Design Checklist: Manual best practices for impactful design, usability, insightfulness, and trustworthiness.
🏅 Dashboards are labeled (Low, High, Golden) based on quality criteria, helping users quickly identify trustworthy insights.
📌 Spotify also developed the Dashboard Portal, a centralized internal platform enhancing dashboard discoverability and providing crucial context like ownership and update frequency.
👉 Read more
🎯 Interested in more on Data Quality scoring and certifications? Airbnb has an excellent deep-dive into their Data Quality Score initiative.
👉 Check out Airbnb’s article
Spotify Engineering
Unlocking Insights with High-Quality Dashboards at Scale | Spotify Engineering
🔥4
🚀 Announcing Next-Level Data Quality Management with DataHub and AI
Exciting news! DataHub is introducing integration with the Model Context Protocol (MCP), unlocking groundbreaking possibilities for data quality management!
Here's what you can expect:
✅ Natural Language Queries: Effortlessly discover metadata by asking simple questions.
✅ Automatic Documentation and Classification: AI-powered agents identify and document your data automatically, significantly reducing your team's workload.
✅ Impact Analysis: Quickly see who and what will be affected by any changes in your data.
✅ IDE Integration: Developers can instantly visualize the impact of data changes directly within their coding environment.
These new capabilities will soon be available both in DataHub Cloud and in open-source, self-hosted setups.
🔗 Learn more about this announcement
📌 Are you excited about incorporating these AI-powered capabilities into your Data Quality processes?
Exciting news! DataHub is introducing integration with the Model Context Protocol (MCP), unlocking groundbreaking possibilities for data quality management!
Here's what you can expect:
✅ Natural Language Queries: Effortlessly discover metadata by asking simple questions.
✅ Automatic Documentation and Classification: AI-powered agents identify and document your data automatically, significantly reducing your team's workload.
✅ Impact Analysis: Quickly see who and what will be affected by any changes in your data.
✅ IDE Integration: Developers can instantly visualize the impact of data changes directly within their coding environment.
These new capabilities will soon be available both in DataHub Cloud and in open-source, self-hosted setups.
🔗 Learn more about this announcement
📌 Are you excited about incorporating these AI-powered capabilities into your Data Quality processes?
pages.acryl.io
Model Context Protocol Support for DataHub
Discover how AI can talk to data with a DataHub MCP server.
💯3✍2
🔍 Deep Dive into dbt docs generate Behavior
📂 Initial Setup:
- manifest.json and catalog.json exist in old_target (previous state)
- No manifest.json or catalog.json yet in the current target
- Modified code for one model: test_source
I ran:
…to better understand how dbt docs generate interacts with the database.
📌 Key Findings:
1. Yes — it runs SQL queries against the database, but only for the selected models (in this case, modified models and their downstream dependencies).
2. Execution Flow:
• 🟦 Project loading & parsing (with partial parsing enabled).
• 🟦 Model compilation — triggers actual SQL queries for selected models and generates manifest.json.
• 🟦 Catalog generation — queries metadata (columns, types, etc.) only for selected models and outputs catalog.json.
3. Confirmed from logs: compilation and catalog steps both involve real DB access — execution time and query results are recorded.
This clarified exactly when and why dbt connects to the database during doc generation — super useful for debugging or performance optimization.
📂 Initial Setup:
- manifest.json and catalog.json exist in old_target (previous state)
- No manifest.json or catalog.json yet in the current target
- Modified code for one model: test_source
I ran:
dbt docs generate --
select=state:modified+ --state=./old_target --debug
…to better understand how dbt docs generate interacts with the database.
📌 Key Findings:
1. Yes — it runs SQL queries against the database, but only for the selected models (in this case, modified models and their downstream dependencies).
2. Execution Flow:
• 🟦 Project loading & parsing (with partial parsing enabled).
• 🟦 Model compilation — triggers actual SQL queries for selected models and generates manifest.json.
• 🟦 Catalog generation — queries metadata (columns, types, etc.) only for selected models and outputs catalog.json.
3. Confirmed from logs: compilation and catalog steps both involve real DB access — execution time and query results are recorded.
This clarified exactly when and why dbt connects to the database during doc generation — super useful for debugging or performance optimization.
✍5
🚀 Continuing the Journey with Model Context Protocol (MCP) and DataHub
In my previous post, I shared the exciting announcement of AI-powered metadata management via MCP in DataHub. Now I want to go deeper.
🧩 MCP for DataHub is already live:
Check out the official repo: github.com/acryldata/mcp-server-datahub
✅ What it does:
• Accepts natural language queries and returns metadata from DataHub
• Supports AI agents and tools (like IDEs, notebooks, or chatbots) to query and act on metadata in a unified way
• Enables semantic context, explanations, and impact-aware actions
🌐 It’s part of a growing ecosystem:
🔹 DBT also introduced its own MCP server: dbt-mcp
Blog post: Introducing dbt-mcp-server
💡 I already tested DataHub’s MCP — it runs smoothly and looks very promising for integrating GenAI with real metadata systems.
📌 Let me know if you’re also exploring MCP or GenAI metadata agents — would love to exchange thoughts!
In my previous post, I shared the exciting announcement of AI-powered metadata management via MCP in DataHub. Now I want to go deeper.
🧩 MCP for DataHub is already live:
Check out the official repo: github.com/acryldata/mcp-server-datahub
✅ What it does:
• Accepts natural language queries and returns metadata from DataHub
• Supports AI agents and tools (like IDEs, notebooks, or chatbots) to query and act on metadata in a unified way
• Enables semantic context, explanations, and impact-aware actions
🌐 It’s part of a growing ecosystem:
🔹 DBT also introduced its own MCP server: dbt-mcp
Blog post: Introducing dbt-mcp-server
💡 I already tested DataHub’s MCP — it runs smoothly and looks very promising for integrating GenAI with real metadata systems.
📌 Let me know if you’re also exploring MCP or GenAI metadata agents — would love to exchange thoughts!
GitHub
GitHub - acryldata/mcp-server-datahub: The official Model Context Protocol (MCP) server for DataHub (https://datahub.com)
The official Model Context Protocol (MCP) server for DataHub (https://datahub.com) - acryldata/mcp-server-datahub
🔥5
🎲 Who is the Best CDO? 🎲
Saw this interesting game shared by Alexandr Barakov from Data Nature and decided to give it a try.
In the game, you step into the shoes of a Chief Data Officer:
- You get emails from colleagues across your company (and occasionally from some sassy cats 🐱) asking for your decisions.
- Your choices influence your budget, data quality, profit, and reputation.
- You'll need to navigate tough choices, messy data, and unexpected cyber-catastrophes.
If you're curious about how you'd handle data leadership challenges, try it out:
👉 whoisthebestcdo.com
Saw this interesting game shared by Alexandr Barakov from Data Nature and decided to give it a try.
In the game, you step into the shoes of a Chief Data Officer:
- You get emails from colleagues across your company (and occasionally from some sassy cats 🐱) asking for your decisions.
- Your choices influence your budget, data quality, profit, and reputation.
- You'll need to navigate tough choices, messy data, and unexpected cyber-catastrophes.
If you're curious about how you'd handle data leadership challenges, try it out:
👉 whoisthebestcdo.com
🔥6
🚨 Soda acquires NannyML! https://launch.soda.io/blog/soda-acquires-nannyml
Soda, leader in data testing and observability, now joins forces with NannyML — open-source library for detecting data drift, concept drift, and silent model failures.
🎯 Why it matters: Soda wants to become the one platform for end-to-end data + ML quality: ✅ Rule-based data checks (pipelines) ✅ ML drift & performance monitoring (production) ✅ No ground truth needed
This move brings Data Engineering and MLOps together on a single platform.
NannyML: https://github.com/NannyML/nannyml
Soda, leader in data testing and observability, now joins forces with NannyML — open-source library for detecting data drift, concept drift, and silent model failures.
🎯 Why it matters: Soda wants to become the one platform for end-to-end data + ML quality: ✅ Rule-based data checks (pipelines) ✅ ML drift & performance monitoring (production) ✅ No ground truth needed
This move brings Data Engineering and MLOps together on a single platform.
NannyML: https://github.com/NannyML/nannyml
GitHub
GitHub - NannyML/nannyml: nannyml: post-deployment data science in python
nannyml: post-deployment data science in python. Contribute to NannyML/nannyml development by creating an account on GitHub.
🔥4👏2
📥 Minimal Permissions for Postgres Ingestion in DataHub
Usually, DataHub needs only metadata, not real data. If you don’t use profiling — reading table rows is not required.
If you want to set permissions safe and precise, here is the minimal working setup:
Tested and works well.
Usually, DataHub needs only metadata, not real data. If you don’t use profiling — reading table rows is not required.
If you want to set permissions safe and precise, here is the minimal working setup:
-- Create user
CREATE USER datahub WITH PASSWORD 'your_strong_password';
-- Access to database and schema
GRANT CONNECT ON DATABASE myproduct TO datahub;
GRANT USAGE ON SCHEMA public TO datahub;
-- Information Schema
GRANT SELECT ON TABLE information_schema.tables TO datahub;
GRANT SELECT ON TABLE information_schema.columns TO datahub;
GRANT SELECT ON TABLE information_schema.views TO datahub;
GRANT SELECT ON TABLE information_schema.schemata TO datahub;
GRANT SELECT ON TABLE information_schema.key_column_usage TO datahub;
GRANT SELECT ON TABLE information_schema.table_constraints TO datahub;
GRANT SELECT ON TABLE information_schema.referential_constraints TO datahub;
GRANT SELECT ON TABLE information_schema.routines TO datahub;
GRANT SELECT ON TABLE information_schema.parameters TO datahub;
-- pg_catalog
GRANT SELECT ON TABLE pg_catalog.pg_class TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_namespace TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_attribute TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_attrdef TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_constraint TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_type TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_enum TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_index TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_inherits TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_depend TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_rewrite TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_proc TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_roles TO datahub;
-- PostgreSQL function
GRANT EXECUTE ON FUNCTION pg_table_size(regclass) TO datahub;
📌 This gives DataHub access to tables, columns, types, constraints, indexes, and more — only metadata.Tested and works well.
✍4🙏1
🚨 New Post
I just published a deep dive into open-source data quality tools — Great Expectations, Soda, Deequ, DQOps, Spark-Expectations, Pandera, and DQX.
What they’re good at, what they’re missing, and where each one fits best
Read here: Comparative Analysis of Open-Source Data Quality Tools
I just published a deep dive into open-source data quality tools — Great Expectations, Soda, Deequ, DQOps, Spark-Expectations, Pandera, and DQX.
What they’re good at, what they’re missing, and where each one fits best
Read here: Comparative Analysis of Open-Source Data Quality Tools
Medium
Overview: Data quality is a critical concern in batch data processing pipelines, as poor-quality…
“” is published by Aleksei Chumagin.
👍3✍2
📦 New version of dag-factory 0.23.0 is out
If you haven’t heard of dag-factory — it’s a tool that lets you create Apache Airflow DAGs from YAML files. No need to write Python code for every DAG. Just define the config and you're done.
In our team, we use dag-factory to generate DAGs that run our Data Quality tests. It works well, configs are simple, and easy to maintain.
I’m also writing this because I contributed to this release:
Updated support for
Added JSON serialization for
Small things, but nice to give back to a tool I use every day.
If you haven’t heard of dag-factory — it’s a tool that lets you create Apache Airflow DAGs from YAML files. No need to write Python code for every DAG. Just define the config and you're done.
In our team, we use dag-factory to generate DAGs that run our Data Quality tests. It works well, configs are simple, and easy to maintain.
I’m also writing this because I contributed to this release:
Updated support for
HttpProvider (PR #389)Added JSON serialization for
HttpOperator (PR #382)Small things, but nice to give back to a tool I use every day.
GitHub
Release v0.23.0 · astronomer/dag-factory
[0.23.0] - 2025-07-14
Breaking Change
Drop Airflow 2.2 Support by @pankajastro in #388
Drop Python 3.8 support by @pankajastro in #435
Drop Airflow 2.3 Support by @pankajastro in #456
Airflow 3 S...
Breaking Change
Drop Airflow 2.2 Support by @pankajastro in #388
Drop Python 3.8 support by @pankajastro in #435
Drop Airflow 2.3 Support by @pankajastro in #456
Airflow 3 S...
👏4👍1