doing something
711 subscribers
910 photos
36 videos
36 files
1.94K links
@smlkw doing something, just my notes to keep updates
Download Telegram
Xnip2025-02-27_16-34-59.png
1.8 MB
How HF space server looks like
doing something
Xnip2025-02-27_16-34-59.png
Understanding Hugging Face platform and their ZeroGPUs

#ai
Media is too big
VIEW IN TELEGRAM
Now my space for Ukrainian Text-to-Speech uses GPU, so the generation is fast.

Use here: https://huggingface.co/spaces/Yehor/radtts-uk-demo

#ai #speech #tts
A new logo in Grok?
Audio
Some examples of different vocoders

#ai #tts #speech
😁2
Forwarded from Serhiy Stetskovych
Коротше я протестував той елевенлабс ASR. Він дуже дуже гарно розпізнає, якщо брати якісну аудіокнигу то там 100% все правильно, але є одна велика проблема. Він не розпізніє літеру Ґ і замість неї постійно пише Г. Також там тейм стемпи тоже дуже добрі але не такі ідеальні і деколи не попападають. На музиці він працює також добре, але слова які не є літературні може не вгадати і таймінги гірші на музиці.
Now I am working on refactoring of RAD-TTS++ inference code to latest torch here:

https://huggingface.co/spaces/Yehor/radtts-uk-vocos-demo

Code is about 2 years old and I wanted to test how it works on MPS device (Apple M cores).

Inference works on MPS, but not all ops are supported:

.../torch/nn/functional.py:5561: UserWarning: The operator 'aten::col2im' is not currently supported on the MPS backend and will fall back to run on the CPU. This may have performance implications. (Triggered internally at /Users/runner/work/pytorch/pytorch/pytorch/aten/src/ATen/mps/MPSFallback.mm:14.)


By the way, RAD-TTS++ and Vocos are super fast on CPU in MacOS.

#ai #torch #ml
doing something
No warnings now
Cleaned the codebase, now it's all you need to do Ukrainian Text-to-Speech
Discovered that justfile supports script execution inside