doing something
https://colab.research.google.com/drive/1o9b2JQ8l9a39uOZZi9DWXQ15BXlCIfEu?usp=sharing 600M model with 4-bits uses about 1GB of VRAM #nlp #ai #hqq
HQQ can be applied to Florence-2 that has an LM at the final layers, as well
https://colab.research.google.com/drive/1G2W0Rxwv-eTe0MJ9aaUuHxLH5lnRc6o_?usp=sharing
In this colab it uses about 1.5 GB to initialize and ~4.5 GB after inference
#ai #cv
https://colab.research.google.com/drive/1G2W0Rxwv-eTe0MJ9aaUuHxLH5lnRc6o_?usp=sharing
In this colab it uses about 1.5 GB to initialize and ~4.5 GB after inference
#ai #cv
Google
Florence-2 + HQQ.ipynb
Colab notebook
doing something
HQQ can be applied to Florence-2 that has an LM at the final layers, as well https://colab.research.google.com/drive/1G2W0Rxwv-eTe0MJ9aaUuHxLH5lnRc6o_?usp=sharing In this colab it uses about 1.5 GB to initialize and ~4.5 GB after inference #ai #cv
I've tested translation of the generated captions using quantized NLLB with longer max_len:
https://colab.research.google.com/drive/1HUgSfm-Y72-1ehorVwtIlzwzXPgrz6YY?usp=sharing
It uses 2 GB of VRAM
#ai #nlp
https://colab.research.google.com/drive/1HUgSfm-Y72-1ehorVwtIlzwzXPgrz6YY?usp=sharing
It uses 2 GB of VRAM
#ai #nlp
Google
Quantized NLLB using HQQ: many sentences.ipynb
Colab notebook
Forwarded from Feed-Master