doing something
Published the testset for Ukrainian to HF: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean #ai #speech
Use the following colabs to see how you can download this dataset in Python:
datasets: https://colab.research.google.com/drive/1qqnr5-WkaJi8iqHa_Pmlx7PbbXwXiimD?usp=sharing
polars: https://colab.research.google.com/drive/1upeXw3WbLjK37b1LetpM0HxFXDdOZqSK?usp=sharing
#ai #asr #speech
datasets: https://colab.research.google.com/drive/1qqnr5-WkaJi8iqHa_Pmlx7PbbXwXiimD?usp=sharing
polars: https://colab.research.google.com/drive/1upeXw3WbLjK37b1LetpM0HxFXDdOZqSK?usp=sharing
#ai #asr #speech
Google
Load Yehor/cv10-uk-testset-clean by datasets.ipynb
Colab notebook
When your work lives own life, but then you notice wonderful things
https://github.com/egorsmkv/qirimtatar-tts-datasets/stargazers
#tts #speech
https://github.com/egorsmkv/qirimtatar-tts-datasets/stargazers
#tts #speech
https://huggingface.co/datasets/Yehor/broadcast-speech-uk
Published a large dataset for ASR in Ukrainian
Dataset is dirty, so need post-processing
#ai #asr #speech
Published a large dataset for ASR in Ukrainian
Dataset is dirty, so need post-processing
#ai #asr #speech
❤1
doing something
Something new in Hugging Face #ai #hf
Hmm, okay it's their DuckDB WASM
> -- The SQL console is powered by DuckDB WASM and runs entirely in the browser.
Hence, it's slow
> -- The SQL console is powered by DuckDB WASM and runs entirely in the browser.
Hence, it's slow
👀3
https://github.com/mit-han-lab/deepcompressor
https://github.com/mit-han-lab/nunchaku
#ai #optimization
https://github.com/mit-han-lab/nunchaku
#ai #optimization
GitHub
GitHub - mit-han-lab/deepcompressor: Model Compression Toolbox for Large Language Models and Diffusion Models
Model Compression Toolbox for Large Language Models and Diffusion Models - mit-han-lab/deepcompressor
A small investigation on WER difference beween two w2v2 models:
https://github.com/egorsmkv/speech-recognition-uk/issues/49
... and why it's important to evaluate models with own dataset.
#ai #speech
https://github.com/egorsmkv/speech-recognition-uk/issues/49
... and why it's important to evaluate models with own dataset.
#ai #speech
GitHub
`Yehor/w2v-bert-uk` vs. `Yehor/w2v-bert-uk-v2.1` · Issue #49 · egorsmkv/speech-recognition-uk
Benchmark table shows that Yehor/w2v-bert-uk is better than Yehor/w2v-bert-uk-v2.1
Published Open Source Crimean Tatar Text-to-Speech dataset to Hugging Face:
https://huggingface.co/datasets/Yehor/qirimtatar-tts
#ai #speech
https://huggingface.co/datasets/Yehor/qirimtatar-tts
#ai #speech
❤2