doing something
https://colab.research.google.com/drive/1o9b2JQ8l9a39uOZZi9DWXQ15BXlCIfEu?usp=sharing 600M model with 4-bits uses about 1GB of VRAM #nlp #ai #hqq
HQQ can be applied to Florence-2 that has an LM at the final layers, as well
https://colab.research.google.com/drive/1G2W0Rxwv-eTe0MJ9aaUuHxLH5lnRc6o_?usp=sharing
In this colab it uses about 1.5 GB to initialize and ~4.5 GB after inference
#ai #cv
https://colab.research.google.com/drive/1G2W0Rxwv-eTe0MJ9aaUuHxLH5lnRc6o_?usp=sharing
In this colab it uses about 1.5 GB to initialize and ~4.5 GB after inference
#ai #cv
Google
Florence-2 + HQQ.ipynb
Colab notebook
doing something
HQQ can be applied to Florence-2 that has an LM at the final layers, as well https://colab.research.google.com/drive/1G2W0Rxwv-eTe0MJ9aaUuHxLH5lnRc6o_?usp=sharing In this colab it uses about 1.5 GB to initialize and ~4.5 GB after inference #ai #cv
I've tested translation of the generated captions using quantized NLLB with longer max_len:
https://colab.research.google.com/drive/1HUgSfm-Y72-1ehorVwtIlzwzXPgrz6YY?usp=sharing
It uses 2 GB of VRAM
#ai #nlp
https://colab.research.google.com/drive/1HUgSfm-Y72-1ehorVwtIlzwzXPgrz6YY?usp=sharing
It uses 2 GB of VRAM
#ai #nlp
Google
Quantized NLLB using HQQ: many sentences.ipynb
Colab notebook
Forwarded from Feed-Master
Phi 3.5 vision + HQQ (4-bits), weights are in bfloat16
vRAM usage decreased: 8.9 GB -> 3.8 GB
https://colab.research.google.com/drive/176YsmMtdy-o0g19HVUPx7gKZ4OQ2tdxl?usp=sharing
#ai #cv #quantization
vRAM usage decreased: 8.9 GB -> 3.8 GB
https://colab.research.google.com/drive/176YsmMtdy-o0g19HVUPx7gKZ4OQ2tdxl?usp=sharing
#ai #cv #quantization
Google
OCR images using Phi-3.5 vision + HQQ (default backend).ipynb
Colab notebook