As always, it's #daily_pain
Converting https://github.com/egorsmkv/cv10-uk-testset-clean/tree/main to HF dataset using sphn
#ai #speech
Converting https://github.com/egorsmkv/cv10-uk-testset-clean/tree/main to HF dataset using sphn
#ai #speech
doing something
As always, it's #daily_pain Converting https://github.com/egorsmkv/cv10-uk-testset-clean/tree/main to HF dataset using sphn #ai #speech
OK, a bug in the data preparation
Should be
Should be
'array': data[0], instead of 'array': data,Published the testset for Ukrainian to HF:
https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean
#ai #speech
https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean
#ai #speech
π₯2
doing something
Published the testset for Ukrainian to HF: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean #ai #speech
Here is an integration of inference code with https://huggingface.co/Yehor/w2v-bert-uk
Colab: https://drive.google.com/file/d/1vBGPGQsLsKI9Yy_MNeageyjK6x-cLehC/view?usp=sharing
#ai #asr #speech
Colab: https://drive.google.com/file/d/1vBGPGQsLsKI9Yy_MNeageyjK6x-cLehC/view?usp=sharing
#ai #asr #speech
https://huggingface.co/Yehor/w2v-xls-r-uk requires about 1.2 GB of GPU memory with float16.
It gives following metrics (without an external LM):
Accuracy on words: 79.76%
Accuracy on chars: 96.36%
Around 300 million of parameters.
#asr #ai #speech
It gives following metrics (without an external LM):
Accuracy on words: 79.76%
Accuracy on chars: 96.36%
Around 300 million of parameters.
#asr #ai #speech
β€1π₯1
doing something
Published the testset for Ukrainian to HF: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean #ai #speech
doing something
What happends when you use Whisper #asr #ai #speech
openai/whisper-large-v3-turbo looks like does not recognize Ρ correctly
it's huggingface checkpoint
it's huggingface checkpoint
β€1
doing something
Evaluation results...
doing something
Published the testset for Ukrainian to HF: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean #ai #speech
Use the following colabs to see how you can download this dataset in Python:
datasets: https://colab.research.google.com/drive/1qqnr5-WkaJi8iqHa_Pmlx7PbbXwXiimD?usp=sharing
polars: https://colab.research.google.com/drive/1upeXw3WbLjK37b1LetpM0HxFXDdOZqSK?usp=sharing
#ai #asr #speech
datasets: https://colab.research.google.com/drive/1qqnr5-WkaJi8iqHa_Pmlx7PbbXwXiimD?usp=sharing
polars: https://colab.research.google.com/drive/1upeXw3WbLjK37b1LetpM0HxFXDdOZqSK?usp=sharing
#ai #asr #speech
Google
Load Yehor/cv10-uk-testset-clean by datasets.ipynb
Colab notebook