π Unlocking High-Quality Dashboards at Scale β Spotify's Approach π§
Spotify shares how they ensure reliable, actionable dashboards for data-driven decision-making at scale. In 2023 alone, Spotify teams created over 4,900 dashboards, actively used by 6,000+ employees.
π Spotifyβs Dashboard Quality Framework:
- Vital Signs: Automated checks ensuring dashboards stay fresh and relevant.
- Spicy Dashboard Design Checklist: Manual best practices for impactful design, usability, insightfulness, and trustworthiness.
π Dashboards are labeled (Low, High, Golden) based on quality criteria, helping users quickly identify trustworthy insights.
π Spotify also developed the Dashboard Portal, a centralized internal platform enhancing dashboard discoverability and providing crucial context like ownership and update frequency.
π Read more
π― Interested in more on Data Quality scoring and certifications? Airbnb has an excellent deep-dive into their Data Quality Score initiative.
π Check out Airbnbβs article
Spotify shares how they ensure reliable, actionable dashboards for data-driven decision-making at scale. In 2023 alone, Spotify teams created over 4,900 dashboards, actively used by 6,000+ employees.
π Spotifyβs Dashboard Quality Framework:
- Vital Signs: Automated checks ensuring dashboards stay fresh and relevant.
- Spicy Dashboard Design Checklist: Manual best practices for impactful design, usability, insightfulness, and trustworthiness.
π Dashboards are labeled (Low, High, Golden) based on quality criteria, helping users quickly identify trustworthy insights.
π Spotify also developed the Dashboard Portal, a centralized internal platform enhancing dashboard discoverability and providing crucial context like ownership and update frequency.
π Read more
π― Interested in more on Data Quality scoring and certifications? Airbnb has an excellent deep-dive into their Data Quality Score initiative.
π Check out Airbnbβs article
Spotify Engineering
Unlocking Insights with High-Quality Dashboards at Scale | Spotify Engineering
π₯4
π Announcing Next-Level Data Quality Management with DataHub and AI
Exciting news! DataHub is introducing integration with the Model Context Protocol (MCP), unlocking groundbreaking possibilities for data quality management!
Here's what you can expect:
β Natural Language Queries: Effortlessly discover metadata by asking simple questions.
β Automatic Documentation and Classification: AI-powered agents identify and document your data automatically, significantly reducing your team's workload.
β Impact Analysis: Quickly see who and what will be affected by any changes in your data.
β IDE Integration: Developers can instantly visualize the impact of data changes directly within their coding environment.
These new capabilities will soon be available both in DataHub Cloud and in open-source, self-hosted setups.
π Learn more about this announcement
π Are you excited about incorporating these AI-powered capabilities into your Data Quality processes?
Exciting news! DataHub is introducing integration with the Model Context Protocol (MCP), unlocking groundbreaking possibilities for data quality management!
Here's what you can expect:
β Natural Language Queries: Effortlessly discover metadata by asking simple questions.
β Automatic Documentation and Classification: AI-powered agents identify and document your data automatically, significantly reducing your team's workload.
β Impact Analysis: Quickly see who and what will be affected by any changes in your data.
β IDE Integration: Developers can instantly visualize the impact of data changes directly within their coding environment.
These new capabilities will soon be available both in DataHub Cloud and in open-source, self-hosted setups.
π Learn more about this announcement
π Are you excited about incorporating these AI-powered capabilities into your Data Quality processes?
pages.acryl.io
Model Context Protocol Support for DataHub
Discover how AI can talk to data with a DataHub MCP server.
π―3β2
π Deep Dive into dbt docs generate Behavior
π Initial Setup:
- manifest.json and catalog.json exist in old_target (previous state)
- No manifest.json or catalog.json yet in the current target
- Modified code for one model: test_source
I ran:
β¦to better understand how dbt docs generate interacts with the database.
π Key Findings:
1. Yes β it runs SQL queries against the database, but only for the selected models (in this case, modified models and their downstream dependencies).
2. Execution Flow:
β’ π¦ Project loading & parsing (with partial parsing enabled).
β’ π¦ Model compilation β triggers actual SQL queries for selected models and generates manifest.json.
β’ π¦ Catalog generation β queries metadata (columns, types, etc.) only for selected models and outputs catalog.json.
3. Confirmed from logs: compilation and catalog steps both involve real DB access β execution time and query results are recorded.
This clarified exactly when and why dbt connects to the database during doc generation β super useful for debugging or performance optimization.
π Initial Setup:
- manifest.json and catalog.json exist in old_target (previous state)
- No manifest.json or catalog.json yet in the current target
- Modified code for one model: test_source
I ran:
dbt docs generate --
select=state:modified+ --state=./old_target --debug
β¦to better understand how dbt docs generate interacts with the database.
π Key Findings:
1. Yes β it runs SQL queries against the database, but only for the selected models (in this case, modified models and their downstream dependencies).
2. Execution Flow:
β’ π¦ Project loading & parsing (with partial parsing enabled).
β’ π¦ Model compilation β triggers actual SQL queries for selected models and generates manifest.json.
β’ π¦ Catalog generation β queries metadata (columns, types, etc.) only for selected models and outputs catalog.json.
3. Confirmed from logs: compilation and catalog steps both involve real DB access β execution time and query results are recorded.
This clarified exactly when and why dbt connects to the database during doc generation β super useful for debugging or performance optimization.
β5
π Continuing the Journey with Model Context Protocol (MCP) and DataHub
In my previous post, I shared the exciting announcement of AI-powered metadata management via MCP in DataHub. Now I want to go deeper.
π§© MCP for DataHub is already live:
Check out the official repo: github.com/acryldata/mcp-server-datahub
β What it does:
β’ Accepts natural language queries and returns metadata from DataHub
β’ Supports AI agents and tools (like IDEs, notebooks, or chatbots) to query and act on metadata in a unified way
β’ Enables semantic context, explanations, and impact-aware actions
π Itβs part of a growing ecosystem:
πΉ DBT also introduced its own MCP server: dbt-mcp
Blog post: Introducing dbt-mcp-server
π‘ I already tested DataHubβs MCP β it runs smoothly and looks very promising for integrating GenAI with real metadata systems.
π Let me know if youβre also exploring MCP or GenAI metadata agents β would love to exchange thoughts!
In my previous post, I shared the exciting announcement of AI-powered metadata management via MCP in DataHub. Now I want to go deeper.
π§© MCP for DataHub is already live:
Check out the official repo: github.com/acryldata/mcp-server-datahub
β What it does:
β’ Accepts natural language queries and returns metadata from DataHub
β’ Supports AI agents and tools (like IDEs, notebooks, or chatbots) to query and act on metadata in a unified way
β’ Enables semantic context, explanations, and impact-aware actions
π Itβs part of a growing ecosystem:
πΉ DBT also introduced its own MCP server: dbt-mcp
Blog post: Introducing dbt-mcp-server
π‘ I already tested DataHubβs MCP β it runs smoothly and looks very promising for integrating GenAI with real metadata systems.
π Let me know if youβre also exploring MCP or GenAI metadata agents β would love to exchange thoughts!
GitHub
GitHub - acryldata/mcp-server-datahub: The official Model Context Protocol (MCP) server for DataHub (https://datahub.com)
The official Model Context Protocol (MCP) server for DataHub (https://datahub.com) - acryldata/mcp-server-datahub
π₯5
π² Who is the Best CDO? π²
Saw this interesting game shared by Alexandr Barakov from Data Nature and decided to give it a try.
In the game, you step into the shoes of a Chief Data Officer:
- You get emails from colleagues across your company (and occasionally from some sassy cats π±) asking for your decisions.
- Your choices influence your budget, data quality, profit, and reputation.
- You'll need to navigate tough choices, messy data, and unexpected cyber-catastrophes.
If you're curious about how you'd handle data leadership challenges, try it out:
π whoisthebestcdo.com
Saw this interesting game shared by Alexandr Barakov from Data Nature and decided to give it a try.
In the game, you step into the shoes of a Chief Data Officer:
- You get emails from colleagues across your company (and occasionally from some sassy cats π±) asking for your decisions.
- Your choices influence your budget, data quality, profit, and reputation.
- You'll need to navigate tough choices, messy data, and unexpected cyber-catastrophes.
If you're curious about how you'd handle data leadership challenges, try it out:
π whoisthebestcdo.com
π₯6
π¨ Soda acquires NannyML! https://launch.soda.io/blog/soda-acquires-nannyml
Soda, leader in data testing and observability, now joins forces with NannyML β open-source library for detecting data drift, concept drift, and silent model failures.
π― Why it matters: Soda wants to become the one platform for end-to-end data + ML quality: β Rule-based data checks (pipelines) β ML drift & performance monitoring (production) β No ground truth needed
This move brings Data Engineering and MLOps together on a single platform.
NannyML: https://github.com/NannyML/nannyml
Soda, leader in data testing and observability, now joins forces with NannyML β open-source library for detecting data drift, concept drift, and silent model failures.
π― Why it matters: Soda wants to become the one platform for end-to-end data + ML quality: β Rule-based data checks (pipelines) β ML drift & performance monitoring (production) β No ground truth needed
This move brings Data Engineering and MLOps together on a single platform.
NannyML: https://github.com/NannyML/nannyml
GitHub
GitHub - NannyML/nannyml: nannyml: post-deployment data science in python
nannyml: post-deployment data science in python. Contribute to NannyML/nannyml development by creating an account on GitHub.
π₯4π2
π₯ Minimal Permissions for Postgres Ingestion in DataHub
Usually, DataHub needs only metadata, not real data. If you donβt use profiling β reading table rows is not required.
If you want to set permissions safe and precise, here is the minimal working setup:
Tested and works well.
Usually, DataHub needs only metadata, not real data. If you donβt use profiling β reading table rows is not required.
If you want to set permissions safe and precise, here is the minimal working setup:
-- Create user
CREATE USER datahub WITH PASSWORD 'your_strong_password';
-- Access to database and schema
GRANT CONNECT ON DATABASE myproduct TO datahub;
GRANT USAGE ON SCHEMA public TO datahub;
-- Information Schema
GRANT SELECT ON TABLE information_schema.tables TO datahub;
GRANT SELECT ON TABLE information_schema.columns TO datahub;
GRANT SELECT ON TABLE information_schema.views TO datahub;
GRANT SELECT ON TABLE information_schema.schemata TO datahub;
GRANT SELECT ON TABLE information_schema.key_column_usage TO datahub;
GRANT SELECT ON TABLE information_schema.table_constraints TO datahub;
GRANT SELECT ON TABLE information_schema.referential_constraints TO datahub;
GRANT SELECT ON TABLE information_schema.routines TO datahub;
GRANT SELECT ON TABLE information_schema.parameters TO datahub;
-- pg_catalog
GRANT SELECT ON TABLE pg_catalog.pg_class TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_namespace TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_attribute TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_attrdef TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_constraint TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_type TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_enum TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_index TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_inherits TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_depend TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_rewrite TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_proc TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_roles TO datahub;
-- PostgreSQL function
GRANT EXECUTE ON FUNCTION pg_table_size(regclass) TO datahub;
π This gives DataHub access to tables, columns, types, constraints, indexes, and more β only metadata.Tested and works well.
β4π1
π¨ New Post
I just published a deep dive into open-source data quality tools β Great Expectations, Soda, Deequ, DQOps, Spark-Expectations, Pandera, and DQX.
What theyβre good at, what theyβre missing, and where each one fits best
Read here: Comparative Analysis of Open-Source Data Quality Tools
I just published a deep dive into open-source data quality tools β Great Expectations, Soda, Deequ, DQOps, Spark-Expectations, Pandera, and DQX.
What theyβre good at, what theyβre missing, and where each one fits best
Read here: Comparative Analysis of Open-Source Data Quality Tools
Medium
Overview: Data quality is a critical concern in batch data processing pipelines, as poor-qualityβ¦
ββ is published by Aleksei Chumagin.
π3β2
π¦ New version of dag-factory 0.23.0 is out
If you havenβt heard of dag-factory β itβs a tool that lets you create Apache Airflow DAGs from YAML files. No need to write Python code for every DAG. Just define the config and you're done.
In our team, we use dag-factory to generate DAGs that run our Data Quality tests. It works well, configs are simple, and easy to maintain.
Iβm also writing this because I contributed to this release:
Updated support for
Added JSON serialization for
Small things, but nice to give back to a tool I use every day.
If you havenβt heard of dag-factory β itβs a tool that lets you create Apache Airflow DAGs from YAML files. No need to write Python code for every DAG. Just define the config and you're done.
In our team, we use dag-factory to generate DAGs that run our Data Quality tests. It works well, configs are simple, and easy to maintain.
Iβm also writing this because I contributed to this release:
Updated support for
HttpProvider (PR #389)Added JSON serialization for
HttpOperator (PR #382)Small things, but nice to give back to a tool I use every day.
GitHub
Release v0.23.0 Β· astronomer/dag-factory
[0.23.0] - 2025-07-14
Breaking Change
Drop Airflow 2.2 Support by @pankajastro in #388
Drop Python 3.8 support by @pankajastro in #435
Drop Airflow 2.3 Support by @pankajastro in #456
Airflow 3 S...
Breaking Change
Drop Airflow 2.2 Support by @pankajastro in #388
Drop Python 3.8 support by @pankajastro in #435
Drop Airflow 2.3 Support by @pankajastro in #456
Airflow 3 S...
π4π1
Tried to connect Trino lineage to DataHub via OpenLineage.
π οΈ Setup: Enabled the OpenLineage Event Listener Sent events to DataHub GMS:
π Root cause: DataHub tries to detect the orchestrator based on the
π Thereβs a pending fix for this: https://github.com/datahub-project/datahub/pull/13066 It adds an explicit check for Trino:
Or route events through a small proxy that rewrites it before forwarding.
β After that β Trino lineage finally shows up in DataHub.
π οΈ Setup: Enabled the OpenLineage Event Listener Sent events to DataHub GMS:
event-listener.name=openlineage
openlineage-event-listener.transport.type=HTTP
openlineage-event-listener.transport.url=https://datahub.internal/openapi/openlineage/api/v1/lineage
openlineage-event-listener.transport.headers=Authorization:Bearer ${ENV:datahub_token}
π First attempt failed β DataHub responded with: java.lang.RuntimeException: Unable to determine orchestrator
π Root cause: DataHub tries to detect the orchestrator based on the
producer field using hardcoded logic. Trino sends: "producer": "https://github.com/trinodb/trino/plugin/trino-openlineage"
β¦but this isnβt recognized β and DataHub crashes with 500.π Thereβs a pending fix for this: https://github.com/datahub-project/datahub/pull/13066 It adds an explicit check for Trino:
else if (producer.startsWith("https://github.com/trinodb/trino/")) {
orchestrator = "trino";
}
π§ͺ Workaround (what worked for me): Just patch the producer field in the JSON to something DataHub understands, like: "producer": "https://github.com/OpenLineage/OpenLineage/blob/v1-0-0/trino"
Or route events through a small proxy that rewrites it before forwarding.
β After that β Trino lineage finally shows up in DataHub.
π3π₯1
I spent 8 hours making Trino usage stats work in DataHub. Here's what I learned.
π§ DataHub ships with
π But what if you're not using Starburst?
Trino supports official event listeners:
- HTTP listener
- Kafka listener
- MySQL event listener β
- OpenLineage listener
I used the MySQL listener. Sounds simple? Not quite.
π§± Problem:
π Solution: You must create a custom VIEW that reshapes raw events into the format
https://gist.github.com/a-chumagin/c9ec0f27d60ca4b10032f0652a1bd034
πͺ« It works β but barely. No full control over email formatting, error handling, or scaling across catalogs.
π§ My recommendation: Use the Trino listener (MySQL, Kafka, etc.), but build your own ingestion in Python with
π₯ Bottom line: DataHub gives you tools β but not always flexibility. If you want usage stats done right, especially outside the Starburst bubble β be ready to write code.
π¦ Iβll be replacing the module with a custom ingestion script β cleaner, portable, and future-proof.
π§ DataHub ships with
starburst-trino-usage β a module designed to pull usage stats (query logs) from a Postgres-based Event Logger."You need to setup Event Logger which saves audit logs into a Postgres db and setup this db as a catalog in Trino."
π But what if you're not using Starburst?
Trino supports official event listeners:
- HTTP listener
- Kafka listener
- MySQL event listener β
- OpenLineage listener
I used the MySQL listener. Sounds simple? Not quite.
π§± Problem:
starburst-trino-usage expects a very specific table schema β but Trino's native listeners log differently.π Solution: You must create a custom VIEW that reshapes raw events into the format
starburst-trino-usage expects. Hereβs the view I used:https://gist.github.com/a-chumagin/c9ec0f27d60ca4b10032f0652a1bd034
πͺ« It works β but barely. No full control over email formatting, error handling, or scaling across catalogs.
π§ My recommendation: Use the Trino listener (MySQL, Kafka, etc.), but build your own ingestion in Python with
MetadataChangeProposalWrapper. This gives you full control β and works not just for Trino, but for any DB that logs events (even Postgres with pgaudit).π₯ Bottom line: DataHub gives you tools β but not always flexibility. If you want usage stats done right, especially outside the Starburst bubble β be ready to write code.
π¦ Iβll be replacing the module with a custom ingestion script β cleaner, portable, and future-proof.
Gist
my_sql_trino.sql
GitHub Gist: instantly share code, notes, and snippets.
π3
π‘ Why DQ is the next big deal
Why Data Quality will soon become important not only in data world.
GenAI is moving crazy fast β new frontier models are coming almost every month. They get better, faster, cheaperβ¦ but there is a limit. Sooner or later, business will start counting money and cut costs. Models will split into two groups:
Work horses β balance between cost and speed
R&D beasts β most expensive and powerful frontier models
And for the βwork horsesβ the key thing will be old principle (GIGO β Garbage In, Garbage Out). If garbage goes in, garbage comes out.
π― Example from my practice We have a RAG service (think retrieval-augmented generation) based on Onyx that answers questions on our internal knowledge base.
Problem: a lot of outdated docs, which means the model ends up answering from old or irrelevant data, and half of them not in English, which means more tokens are spent for translation or handling that extra text. No matter how we tune the prompt, it only gets bigger (means more expensive), and keeping it up to date is pain.
Same story with most AI agents. For demo, to impress CEO β perfect. But in production you quickly see: the data they use is the real boss here.
And this is not just my observation β research points exactly the same direction.
π Not only us: In Journal of Scientific and Engineering Research link they say:
βAI is determined by the quality of data input that feeds its modelsβ¦ AI models created from low-quality or predominantly biased or incomplete information will produce distortions.β βData is the lifeblood of any AI model, and the quality of data that feeds into an AI model determines the capability of the resulting model, its precision, stability, and equity.β
And yes, hello from Data Governance β companies want to manage not only quality but also access to data that models can touch.
π― Final thought You can replace a model in one day. But cleaning and organizing data β itβs long, boring, and hard work. And itβs the thing that decides if your AI will be a growth tool or an error generator.
β¦And believe me, no any prompt will clean that data for you.
Why Data Quality will soon become important not only in data world.
GenAI is moving crazy fast β new frontier models are coming almost every month. They get better, faster, cheaperβ¦ but there is a limit. Sooner or later, business will start counting money and cut costs. Models will split into two groups:
Work horses β balance between cost and speed
R&D beasts β most expensive and powerful frontier models
And for the βwork horsesβ the key thing will be old principle (GIGO β Garbage In, Garbage Out). If garbage goes in, garbage comes out.
π― Example from my practice We have a RAG service (think retrieval-augmented generation) based on Onyx that answers questions on our internal knowledge base.
Problem: a lot of outdated docs, which means the model ends up answering from old or irrelevant data, and half of them not in English, which means more tokens are spent for translation or handling that extra text. No matter how we tune the prompt, it only gets bigger (means more expensive), and keeping it up to date is pain.
Same story with most AI agents. For demo, to impress CEO β perfect. But in production you quickly see: the data they use is the real boss here.
And this is not just my observation β research points exactly the same direction.
π Not only us: In Journal of Scientific and Engineering Research link they say:
βAI is determined by the quality of data input that feeds its modelsβ¦ AI models created from low-quality or predominantly biased or incomplete information will produce distortions.β βData is the lifeblood of any AI model, and the quality of data that feeds into an AI model determines the capability of the resulting model, its precision, stability, and equity.β
And yes, hello from Data Governance β companies want to manage not only quality but also access to data that models can touch.
π― Final thought You can replace a model in one day. But cleaning and organizing data β itβs long, boring, and hard work. And itβs the thing that decides if your AI will be a growth tool or an error generator.
β¦And believe me, no any prompt will clean that data for you.
π―3β€2
Case study: DataHub βMonthly Active Usersβ β unique people
While digging into the DataHub code, I found that the MAU highlight in the UI does not count unique users.
It counts unique browsers.
- A single person can appear multiple times if they use different devices, incognito, or clear cookies.
- Thatβs why MAU can be noticeably higher than your actual unique users.
---
How itβs calculated
The GraphQL highlight counts the cardinality of `browserId` over the last month, excluding backend events, using timestamp (event time).
Code proof:
---
If you want real unique users β query by actorUrn instead of browserId.
All examples below target:
/openapi/v2/analytics/datahub_usage_events/_search
---
1. Unique users (front-end only)
---
2. Browsers per user
---
3. Replicate MAU exactly (for validation)
---
Notes
- Always use timestamp (event time) instead of @ timestamp (ingest time) to avoid backfill/reindex noise.
- Exclude usageSource=backend to focus on UI-generated activity.
Key takeaways
- DataHub MAU counts unique browsers, not unique people.
- Expect MAU > unique users if the same person uses multiple browsers or devices.
- Calculate uniq user with request to datahub_usage_events
While digging into the DataHub code, I found that the MAU highlight in the UI does not count unique users.
It counts unique browsers.
- A single person can appear multiple times if they use different devices, incognito, or clear cookies.
- Thatβs why MAU can be noticeably higher than your actual unique users.
---
How itβs calculated
The GraphQL highlight counts the cardinality of `browserId` over the last month, excluding backend events, using timestamp (event time).
Code proof:
int activeUsersThisRange =
_analyticsService.getHighlights(
_analyticsService.getUsageIndexName(),
Optional.of(dateRangeThis),
ImmutableMap.of(),
ImmutableMap.of(),
Optional.of("browserId"));
public String getUsageIndexName() {
return _indexConvention.getIndexName(DATAHUB_USAGE_EVENT_INDEX);
}
private QueryBuilder getDefaultFilters() {
return QueryBuilders.boolQuery()
.mustNot(
QueryBuilders.termQuery(
DataHubUsageEventConstants.USAGE_SOURCE,
DataHubUsageEventConstants.BACKEND_SOURCE));
}
private AggregationBuilder getFilteredAggregation(
Map<String, List<String>> mustFilters,
Map<String, List<String>> mustNotFilters,
Optional<DateRange> dateRange) {
// Use timestamp as dateRangeField
return getFilteredAggregation(mustFilters, mustNotFilters, dateRange, "timestamp");
}
---
If you want real unique users β query by actorUrn instead of browserId.
All examples below target:
/openapi/v2/analytics/datahub_usage_events/_search
---
1. Unique users (front-end only)
{
"query": {
"bool": {
"must": [
{ "range": { "timestamp": { "gte": "2025-07-13T00:00:00Z", "lt": "2025-08-14T00:00:00Z" } } }
],
"must_not": [
{ "term": { "usageSource": "backend" } }
]
}
},
"aggs": {
"unique_actors": {
"cardinality": {
"field": "actorUrn.keyword",
"precision_threshold": 40000
}
}
},
"size": 0
}
---
2. Browsers per user
{
"query": {
"bool": {
"must": [
{ "range": { "timestamp": { "gte": "2025-07-13T00:00:00Z", "lt": "2025-08-14T00:00:00Z" } } }
],
"must_not": [
{ "term": { "usageSource": "backend" } }
]
}
},
"aggs": {
"actors": {
"terms": { "field": "actorUrn.keyword", "size": 10000 },
"aggs": {
"unique_browsers": {
"cardinality": {
"field": "browserId",
"precision_threshold": 40000
}
}
}
}
},
"size": 0
}
---
3. Replicate MAU exactly (for validation)
{
"query": {
"bool": {
"must": [
{ "range": { "timestamp": { "gte": "2025-07-13T00:00:00Z", "lt": "2025-08-14T00:00:00Z" } } }
],
"must_not": [
{ "term": { "usageSource": "backend" } }
]
}
},
"aggs": {
"mau": {
"cardinality": {
"field": "browserId",
"precision_threshold": 40000
}
}
},
"size": 0
}
---
Notes
- Always use timestamp (event time) instead of @ timestamp (ingest time) to avoid backfill/reindex noise.
- Exclude usageSource=backend to focus on UI-generated activity.
Key takeaways
- DataHub MAU counts unique browsers, not unique people.
- Expect MAU > unique users if the same person uses multiple browsers or devices.
- Calculate uniq user with request to datahub_usage_events
π₯4