Tested Intel's quantized models of Whisper made by Neural Compressor:
https://github.com/egorsmkv/whisper-intel-optimized
#ai #asr #quantization
https://github.com/egorsmkv/whisper-intel-optimized
#ai #asr #quantization
GitHub
GitHub - egorsmkv/optimized-whisper-intel: Run quantized Whisper models only on CPU with Intel hardware
Run quantized Whisper models only on CPU with Intel hardware - egorsmkv/optimized-whisper-intel
Made first Streamlit app (converting Gradio app using GitHub Copilot)
https://huggingface.co/spaces/Yehor/st-hubert-uk-demo
#ai #ui
https://huggingface.co/spaces/Yehor/st-hubert-uk-demo
#ai #ui
huggingface.co
Streamlit Speech-to-Text for Ukrainian (HuBERT) - a Hugging Face Space by Yehor
Discover amazing ML apps made by the community
doing something
https://mobiusml.github.io/whisper-static-cache-blog/ #asr #ai #whisper #quantization
GitHub
RuntimeError: Expected in.dtype() == at::kInt to be true, but got false. · Issue #108 · mobiusml/hqq
I want to reproduce Whisper + HQQ example but getting error: --------------------------------------------------------------------------- RuntimeError Traceback (most recent call last) [<ipython-...
Okay, I've made a working colab with these speeding mechanisms for Whisper:
- HQQ (quantize linear layers to 4-bit)
- torch.compile
- use static cache
Colab:
https://colab.research.google.com/drive/1zERp53F7gmE58OiqFR0iNV3gw3ybiwHM?usp=sharing
RTF: 0.0531 (batch size = 1) with large-v3
#ai #whisper #asr
- HQQ (quantize linear layers to 4-bit)
- torch.compile
- use static cache
Colab:
https://colab.research.google.com/drive/1zERp53F7gmE58OiqFR0iNV3gw3ybiwHM?usp=sharing
RTF: 0.0531 (batch size = 1) with large-v3
#ai #whisper #asr
Google
Nightly torch with HQQ (nbits=4) + torch.compile + static cache + Whisper.ipynb
Colab notebook
Forwarded from Feed-Master
rerun.io
Why Rust?
I've been a programmer for 20+ years, and few things excite me as much as Rust. My background is mostly in C++, though I have also worked in Python and Lua, and dabbled in many more languages. I started writing Rust around 2014, and since 2018 I've been writing…
doing something
bs=24, whisper large-v3 https://github.com/egorsmkv/optimized-whisper #ai #asr
Forwarded from Feed-Master
burn.dev
Burn 0.14.0 Release Notes
This release marks the debut of our CubeCL integration, which brings cross-platform GPU programming capabilities directly to Rust.
As always, it also includes numerous bug fixes, performance enhancements, new tensor operations, and improved documentation.
As always, it also includes numerous bug fixes, performance enhancements, new tensor operations, and improved documentation.