Made first Streamlit app (converting Gradio app using GitHub Copilot)
https://huggingface.co/spaces/Yehor/st-hubert-uk-demo
#ai #ui
https://huggingface.co/spaces/Yehor/st-hubert-uk-demo
#ai #ui
huggingface.co
Streamlit Speech-to-Text for Ukrainian (HuBERT) - a Hugging Face Space by Yehor
Discover amazing ML apps made by the community
doing something
https://mobiusml.github.io/whisper-static-cache-blog/ #asr #ai #whisper #quantization
GitHub
RuntimeError: Expected in.dtype() == at::kInt to be true, but got false. · Issue #108 · mobiusml/hqq
I want to reproduce Whisper + HQQ example but getting error: --------------------------------------------------------------------------- RuntimeError Traceback (most recent call last) [<ipython-...
Okay, I've made a working colab with these speeding mechanisms for Whisper:
- HQQ (quantize linear layers to 4-bit)
- torch.compile
- use static cache
Colab:
https://colab.research.google.com/drive/1zERp53F7gmE58OiqFR0iNV3gw3ybiwHM?usp=sharing
RTF: 0.0531 (batch size = 1) with large-v3
#ai #whisper #asr
- HQQ (quantize linear layers to 4-bit)
- torch.compile
- use static cache
Colab:
https://colab.research.google.com/drive/1zERp53F7gmE58OiqFR0iNV3gw3ybiwHM?usp=sharing
RTF: 0.0531 (batch size = 1) with large-v3
#ai #whisper #asr
Google
Nightly torch with HQQ (nbits=4) + torch.compile + static cache + Whisper.ipynb
Colab notebook
Forwarded from Feed-Master
rerun.io
Why Rust?
I've been a programmer for 20+ years, and few things excite me as much as Rust. My background is mostly in C++, though I have also worked in Python and Lua, and dabbled in many more languages. I started writing Rust around 2014, and since 2018 I've been writing…
doing something
bs=24, whisper large-v3 https://github.com/egorsmkv/optimized-whisper #ai #asr
Forwarded from Feed-Master
burn.dev
Burn 0.14.0 Release Notes
This release marks the debut of our CubeCL integration, which brings cross-platform GPU programming capabilities directly to Rust.
As always, it also includes numerous bug fixes, performance enhancements, new tensor operations, and improved documentation.
As always, it also includes numerous bug fixes, performance enhancements, new tensor operations, and improved documentation.
Fast denoiser based on diffusions:
https://github.com/sp-uhh/sgmse_crp
RTF about 0.0375 on RTX 4k Ada, because it needs only 1 reverse step
#ai #diffusions #speech
https://github.com/sp-uhh/sgmse_crp
RTF about 0.0375 on RTX 4k Ada, because it needs only 1 reverse step
#ai #diffusions #speech
GitHub
GitHub - sp-uhh/sgmse_crp
Contribute to sp-uhh/sgmse_crp development by creating an account on GitHub.
We’ve been talking about PDF processing yesterday at UDS and today I’ve discovered it:
https://x.com/hu_yifei/status/1828870309857915341?s=46&t=7jwH29MvU0R301CgvqVBYw
#cv #ai
https://x.com/hu_yifei/status/1828870309857915341?s=46&t=7jwH29MvU0R301CgvqVBYw
#cv #ai
This author has another model (based on Florence-2) that extracts parts in the article: https://huggingface.co/yifeihu/TFT-ID-1.0
In the album some examples how it works. Next we can OCR these images and analyse using other LLMs.
#ai #cv
In the album some examples how it works. Next we can OCR these images and analyse using other LLMs.
#ai #cv