Dateno
105 subscribers
6 photos
16 links
Discover, search, integrate: Open data at your fingertips. More about Dateno https://t.ly/z8v7g
Download Telegram
Channel created
Channel photo updated
A few months ago we launched Dateno, a new search engine with many unique features that we are proud of.

Firstly, Dateno is a focused search engine, similar to many academic search engines or Google Dataset Search.

At it’s core is the Common Data Index core, now renamed Dateno Registry. It’s more than 10,000 data catalogues all over the world, with almost every data catalogue linked to the country, certain topics and so on. For about a year we have been collecting these data catalogues using various discovery methods.

At some point, this registry became so large that it’s possible to create the metadata crawlers that collect details about all possible datasets.

Read more about Dateno at https://t.ly/z8v7g
🔥2❤‍🔥1❤1
Exciting News from Dateno!

We are thrilled to announce that Dateno has successfully closed its latest investment round, led by Blockchair! 🎉 This marks a major milestone in our mission to revolutionize data accessibility and search.

Since our launch just a few months ago, Dateno has been rapidly growing, now indexing over 15 million datasets. By the end of 2024, we aim to expand this number to 30 million! Our platform offers a focused and advanced data search experience, supporting 13 facets for filtering results, making it easier than ever for users to find the datasets they need.
With this new investment and partnership, we’re excited to roll out major updates, including the launch of the Dateno API. This will position Dateno as the world's largest search index for data, allowing other projects to integrate our robust data search capabilities directly into their platforms.

We’re also incorporating blockchain and web3 data from Blockchair and other decentralized finance players, and we’re hard at work on AI-powered features to improve search accuracy and relevance. These enhancements will empower data analysts worldwide, making their work more intuitive, efficient, and insightful.

We’re just getting started, and we’re grateful for the support of our investors, partners, and the entire Dateno community. Stay tuned for more updates, and thank you for being part of this journey with us! 🚀✨

#Dateno #DataSearch #Investment #Innovation #AI
🍾8🎉3❤2
Dateno Expands Data Capabilities for Professionals with API and Dashboard Tools!

We are thrilled to announce the launch of two powerful tools designed specifically for data professionals: the My Dateno personal dashboard and the Dateno API! These updates will greatly enhance your ability to manage and integrate data search into your workflows.

With My Dateno, users can now track their search history and access API keys, making it easier than ever to tap into Dateno's extensive data search capabilities. In the future, My Dateno will also provide access to premium features and additional data services. Plus, those who join our early access program will get free access to these new features during the testing period!

The Dateno API enables developers and businesses to integrate our platform’s search functionality directly into their products and infrastructure. This API offers fast, efficient search across 19 million datasets—including data files, geoAPI connections, and statistical indicators—with powerful filtering options. Retrieve comprehensive metadata and related resources, and streamline your data processing with ease.

We’re excited to empower data professionals with these new tools! 🚀

Learn more and sign up for early access at dateno.io

#Dateno #DataSearch #API #Innovation #DataIntegration #DataProfessionals
🔥9
🚀 Dateno Enters Industrial Operation – Redefining Global Dataset Search
We’re excited to announce that Dateno has officially transitioned to full-scale industrial operation! 🎉 Now, data professionals worldwide can seamlessly access over 20 million high-quality datasets with advanced filtering, API integration, and continuously updated sources.

🔍 What makes Dateno stand out?
✅ Extensive dataset collection – 20M+ datasets indexed, aiming for 30M.
✅ Advanced filtering – Search by dataset owner, geography, topic, and more.
✅ AI-powered search – Recognizes semantic relationships (DOI, geolocation).
✅ API-first approach – Seamless integration into analytics & ML pipelines.
✅ High-quality, ad-free data – Focused on clean, structured, and trustworthy datasets.

💡 What’s next?
🔹 Expanding the dataset index to cover even more industries & research fields.
🔹 Improving search quality & user experience.
🔹 Enhancing AI-driven search for more relevant results.
🔹 Adding new API capabilities for seamless integration.
🔹 Launching tools to help professionals derive deeper insights.

Dateno is more than a search engine – it’s an ecosystem built to make data discovery effortless. 🌍

Join us and experience the next level of fast, precise, and integrated dataset search!
👉 Learn more: dateno.io
📩 Contact us: dateno@dateno.io

#Dateno #DataSearch #MachineLearning #BigData #AI
🔥6👍4⚡3
Global stats just got a major upgrade at Dateno!

We’ve updated time series from the World Bank (DataBank) and International Labour Organization (ILOSTAT) — now available in a more powerful and usable format.

📊 What’s new?
19,000+ indicators across economics, employment, trade, health & more
3.85 million time series with clean structure and rich metadata
Support for multiple export formats: CSV, Excel, JSON, Stata, Parquet, and more
Fully documented schemas and all source metadata included
We’re not just expanding our data coverage — we’re raising the bar for how usable and reliable open statistical data can be.

And there’s more coming:
📡 New sources of global indicators
🧠 Improved dataset descriptions
🧩 A specialized API for working with time series in extended formats
Have a specific use case for international statistics? We’d love to hear from you → dateno@dateno.io

🔍 Try it now: https://dateno.io

#openData #datadiscovery #statistics #dataengineering #dateno #worldbank #ILOSTAT
🔥5❤2👍1
🔍 One of the key features of the Dateno search engine is that, in addition to collecting basic metadata about datasets and APIs, its crawlers also gather links to related resources and even archive some of them.

This approach not only helps provide users with a convenient tool for finding data but also allows us to analyze how data is actually published and in what formats.

📊 As of July 2025, Dateno has indexed 5,961,849 datasets from open data portals. That’s about 27% of the total datasets, map layers, and time series aggregated from data catalogs, geoportals, and statistical databases.

Let’s take a closer look at these 5.9 million datasets:
Some datasets come without any associated files, while others may include dozens or even hundreds of attached resources. That’s why, when analyzing file formats, it makes more sense to focus on the number of resources rather than the number of datasets.

👉 Currently, Dateno indexes 6.7 million resources (files and links) attached to these datasets—on average, around 1.1 resources per dataset.

Here’s the breakdown of the most common file formats:

CSV: 1,008,646 files (15%)

XLSX: 525,329 files (7.8%)

XML: 522,501 files (7.8%)

JSON: 509,668 files (7.6%)

ZIP: 496,709 files (7.4%)

PDF: 487,189 files (7.3%)

HTML: 475,377 files (7.1%)

WMS (geospatial API): 320,159 files (4.8%)

NC (NetCDF): 233,229 files (3.5%)

XLS: 185,855 files (2.8%)

WCS (geospatial API): 141,472 files (2.1%)

KML: 122,781 files (1.8%)

DOCX: 115,723 files (1.7%)
…and many more.

It’s no surprise that CSV remains the most popular format for open data publication. Other common formats include XLSX, XML, JSON, and legacy XLS files.

Formats like WCS, WMS, and KML reflect the increasing role of geospatial data published via standardized APIs and file formats.

Meanwhile, the popularity of PDF, DOCX, and HTML points to the reality that not all datasets come with machine-readable files. Sometimes data is shared as reports, documents, or links to external sources, requiring additional effort to extract the actual data.

📉 And what about data science-friendly formats?
Take Parquet files, for example—widely used in data engineering and data science for their efficiency. Surprisingly, only 1,652 Parquet files are currently indexed by Dateno—less than 0.025% of all resources. Quite an eye-opener!

The world of open data is still far from being fully aligned with the needs of data engineering and data science. Closing this gap is essential if we want to unlock the full potential of open data for advanced analytics and AI.

#OpenData #DataScience #DataEngineering #Dateno #DataFormats #Parquet #CSV #GeospatialData #AI
✍3🔥3❤1
Almost 2.5 years ago, I wrote a longread (in Russian) about Uzbekistan’s open data portal—“What’s Wrong with Uzbekistan’s Open Data Portal?”. Recently, I decided to take another look at the portal: data.egov.uz. Unfortunately, not much has changed.

Yes, the number of datasets has grown—from 6,623 to 10,412. That sounds impressive. But here’s the reality:

➡️ In 2023, there were 2,823 single-row datasets. Today, there are 5,207 of them—50% of the entire portal.
➡️ Only 114 datasets contain more than 1,000 records—that's just over 1% of all published datasets.
➡️ The total uncompressed volume of the portal’s data (in JSON) is around 426 MB (compared to 284 MB last year).

Why create thousands of datasets that each contain just a single row? The answer is simple: to boost quantity, not quality. For real data users, such datasets are virtually useless.

Has anything meaningfully changed with open data in Uzbekistan? Sadly, no. The number of datasets is not a true indicator of openness when the majority are artificially fragmented like this.

At Dateno (dateno.io), we’ve chosen not to index this portal—at least for now. It uses non-standard software, isn’t easily crawlable, and more than half of its datasets lack meaningful value.

👉 What do you think?
Do such open data portals provide any real value?
Is it worth talking about them at all?

#OpenData #Transparency #DataQuality #Dateno #DigitalGovernance #DataPortals
🔥4
🚀 Russia’s Open Data Portal Relaunch – But With a Twist

The Russian Ministry of Economic Development has relaunched data.gov.ru after a 2‑year shutdown.

🔎 What’s new?

The portal now hosts ~5,000 datasets, compared to 24,000+ datasets back in early 2022.

The total size of compressed data is about 100 MB, versus 14 GB in 2022.

Most datasets are small CSV files that haven’t been updated in 4–10 years.

📉 What’s missing?
Before 2022, the portal was often called a “data dump” full of outdated information — but at least there was a lot of it. Now, despite the relaunch, the volume and freshness of data have dramatically decreased.

💡 Our next step:
We’ve made a full dump of the new portal and are evaluating whether it’s worth indexing in our search engine Dateno. Early signs: unfortunately, there’s not much of real value there.

📌 I’ll share a more detailed analysis of the portal’s contents and structure soon. Stay tuned!

hashtag#opendata hashtag#russia hashtag#dataportal
✍2👏2👍1
Exploring Australia’s Research Data Portal – researchdata.edu.au

Did you know that Australia has a dedicated Research Data Portal that aggregates an impressive 224,000 datasets, with 96,000 of them available online?

This portal works as a powerful search engine across dozens of academic repositories, archives, government open data portals, and geospatial portals. In many ways, it feels similar to Dateno, offering search across nine types of facets (filters).

What’s even more interesting is that it doesn’t stop at datasets: you can also search for research projects, people and organizations, services, and software products, among others. A large share of the materials are published under open licenses.

For comparison, Dateno currently lists 676,000 datasets related to Australia, mostly from open data and geospatial portals. However, it includes far fewer research datasets — largely because there are already strong specialized tools like this portal. In that sense, Research Data and Dateno complement each other rather than compete.

One note: the Research Data portal has very few statistical datasets or time series, which is surprising given Australia’s advanced official statistics publishing systems.

📌 We likely won’t index this portal directly in Dateno — but indexing the original sources that feed into it is definitely on our radar.

👉 If you work with research data or open data in Australia, this is a resource worth bookmarking!
❤1
Data engineer needed!

We are looking for a data engineer to develop an ambitious modern dataset search engine Dateno (dateno.io). Fully remote

Today the technology stack includes FastAPI, Airflow, MongoDB, Elasticsearch. We use Github + Discord for management.

Our technology stack more https://stackshare.io/dateno/dateno

Responsibilities:
Development and maintaining of Dateno data infrastructure
Preparing, adjusting and monitoring data pipelines
Resolving data quality issues

Requirements:
Experience with Python data stack 1+ year with real product;
Experience with building data pipelines with open source data stack;
Understating data quality management and monitoring;
Knowledge of the data observability issues and frameworks
Experience with REST API;
Knowledge of English at the level of reading technical documentation and basic communication;
Strong technical problem solving skills
Responsibility, ability to work independently.

Pros are:
Data engineering education: MS degree or equivalent industry experience
Experience or willingness to work with NoSQL databases such as MongoDB and Elasticsearch;
Experience and willingness to use modern database engines stack as DuckDB, Clickhouse and e.t.c.
Portfolio - github link with example projects/modules/code/contributions to open source projects;
Love for open data and open source is a definite plus.

Conditions: Full-time, salary based on the results of the interview.

The main thing - compliance with deadlines and the desire to make the world a better place.

Company: Dateno
Contact: dateno@dateno.io
⚡3