https://colab.research.google.com/drive/1o9b2JQ8l9a39uOZZi9DWXQ15BXlCIfEu?usp=sharing
600M model with 4-bits uses about 1GB of VRAM
#nlp #ai #hqq
600M model with 4-bits uses about 1GB of VRAM
#nlp #ai #hqq
Google
Quantized NLLB using HQQ.ipynb
Colab notebook
doing something
decord can lag (it does not get the batch of frames) on GPU device if they have high values
This story has ended with a script that uses StreamReader from torchaudio.
I'm using the seek method to skip unrelated frames.
One important thing is that you need compiled ffmpeg with CUDA to do appropriate decoding.
I'm using the seek method to skip unrelated frames.
One important thing is that you need compiled ffmpeg with CUDA to do appropriate decoding.
An idea for a weekend project:
Adapt https://github.com/Gadersd/whisper-burn to the recently publish Whisper Turbo model.
Need to make only two things:
- update the repo to latest burn 🔥 version
- get down the number of decoder layers
#ai #asr
Adapt https://github.com/Gadersd/whisper-burn to the recently publish Whisper Turbo model.
Need to make only two things:
- update the repo to latest burn 🔥 version
- get down the number of decoder layers
#ai #asr
GitHub
GitHub - Gadersd/whisper-burn: A Rust implementation of OpenAI's Whisper model using the burn framework
A Rust implementation of OpenAI's Whisper model using the burn framework - Gadersd/whisper-burn
https://youtube.com/playlist?list=PLoROMvodv4rPOWA-omMM6STXaWW4FvJT8&feature=shared
https://deepgenerativemodels.github.io/
#ai #genai
https://deepgenerativemodels.github.io/
#ai #genai
YouTube
Stanford CS236: Deep Generative Models I 2023 I Stefano Ermon
For more information about Stanford's Artificial Intelligence programs visit: https://stanford.io/ai View the course website: https://deepgenerativemodels.gi...
Finally we can cut texts from memes to make other ones
https://huggingface.co/spaces/OzzyGT/diffusers-image-fill
#ai #genai
https://huggingface.co/spaces/OzzyGT/diffusers-image-fill
#ai #genai
doing something
https://colab.research.google.com/drive/1o9b2JQ8l9a39uOZZi9DWXQ15BXlCIfEu?usp=sharing 600M model with 4-bits uses about 1GB of VRAM #nlp #ai #hqq
HQQ can be applied to Florence-2 that has an LM at the final layers, as well
https://colab.research.google.com/drive/1G2W0Rxwv-eTe0MJ9aaUuHxLH5lnRc6o_?usp=sharing
In this colab it uses about 1.5 GB to initialize and ~4.5 GB after inference
#ai #cv
https://colab.research.google.com/drive/1G2W0Rxwv-eTe0MJ9aaUuHxLH5lnRc6o_?usp=sharing
In this colab it uses about 1.5 GB to initialize and ~4.5 GB after inference
#ai #cv
Google
Florence-2 + HQQ.ipynb
Colab notebook
doing something
HQQ can be applied to Florence-2 that has an LM at the final layers, as well https://colab.research.google.com/drive/1G2W0Rxwv-eTe0MJ9aaUuHxLH5lnRc6o_?usp=sharing In this colab it uses about 1.5 GB to initialize and ~4.5 GB after inference #ai #cv
I've tested translation of the generated captions using quantized NLLB with longer max_len:
https://colab.research.google.com/drive/1HUgSfm-Y72-1ehorVwtIlzwzXPgrz6YY?usp=sharing
It uses 2 GB of VRAM
#ai #nlp
https://colab.research.google.com/drive/1HUgSfm-Y72-1ehorVwtIlzwzXPgrz6YY?usp=sharing
It uses 2 GB of VRAM
#ai #nlp
Google
Quantized NLLB using HQQ: many sentences.ipynb
Colab notebook