Forwarded from Задуха
З допомогою @smlkw була втілена в життя моя давня ідея — створити каталог всіх українських телеграм ресурсів які мають відношення до ШІ.
Дуже радію що нарешті покладений початок створенню мережі, в яку вже війшла велика кількість ресурсів (про багато з яких я сам дізнався лише сьогодні). Впевненний це допоможе створювати більше горизонтальних зв'язків та нових проектів
https://t.me/nlp_uk/15645
Дуже радію що нарешті покладений початок створенню мережі, в яку вже війшла велика кількість ресурсів (про багато з яких я сам дізнався лише сьогодні). Впевненний це допоможе створювати більше горизонтальних зв'язків та нових проектів
https://t.me/nlp_uk/15645
Telegram
Yehor Smoliakov in NLP української мови
Шановні мембери цього чату, ми тут з Богданом організували збір каналів та груп де пишуть та спілкуються українською про ШІ:
https://t.me/addlist/3l-AohadWpcwMzBi
Наразі зібрали 30 штук, але впевнені — це не повний список, тому просимо вашої допомоги. …
https://t.me/addlist/3l-AohadWpcwMzBi
Наразі зібрали 30 штук, але впевнені — це не повний список, тому просимо вашої допомоги. …
doing something
https://github.com/paradedb/paradedb #rust
Yesterday I’ve integrated their PgSQL extension with BM25 to my software and it works awesome
What it is:
https://en.m.wikipedia.org/wiki/Okapi_BM25
What it is:
https://en.m.wikipedia.org/wiki/Okapi_BM25
Wikipedia
Okapi BM25
ranking function used by search engines
Forwarded from Hacker News
I highly recommend to use Typst as an alternative to write papers, instead of LaTeX.
Many authors send them to arxiv.org so we need a base template to start with, so we have one:
https://github.com/mgoulao/arkheion
#papers
Many authors send them to arxiv.org so we need a base template to start with, so we have one:
https://github.com/mgoulao/arkheion
#papers
GitHub
GitHub - mgoulao/arkheion: A Typst template for arXiv
A Typst template for arXiv. Contribute to mgoulao/arkheion development by creating an account on GitHub.
How to convert PDFs to images:
https://colab.research.google.com/drive/1Rykfvul5mwE-bRSkQAEDxMLRdWjSAiH1?usp=sharing
Frequent task in document processing, so let's it be here
#python
https://colab.research.google.com/drive/1Rykfvul5mwE-bRSkQAEDxMLRdWjSAiH1?usp=sharing
Frequent task in document processing, so let's it be here
#python
Google
pdf2image.ipynb
Colab notebook
doing something
How to convert PDFs to images: https://colab.research.google.com/drive/1Rykfvul5mwE-bRSkQAEDxMLRdWjSAiH1?usp=sharing Frequent task in document processing, so let's it be here #python
Made a version of the same with MuPDF:
https://colab.research.google.com/drive/1OioZY1vvMVHLx3F4XWKF848moZtuQUCw?usp=sharing
Thanks to @junior_programer for his notes
https://colab.research.google.com/drive/1OioZY1vvMVHLx3F4XWKF848moZtuQUCw?usp=sharing
Thanks to @junior_programer for his notes
Google
PyMuPDF: convert PDF to Images.ipynb
Colab notebook
An idea for your resume project or how to make a project with CTO’s excitement during interview:
Here’s a large dataset: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech
You can create an ETL pipeline to prepare the data for training an ASR model.
With parallelization and fast audio resampling, you can use Apache Airflow to solve the task.
#thoughts
Here’s a large dataset: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech
You can create an ETL pipeline to prepare the data for training an ASR model.
With parallelization and fast audio resampling, you can use Apache Airflow to solve the task.
#thoughts
huggingface.co
MLCommons/unsupervised_peoples_speech · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.