Forwarded from Feed-Master
Phi 3.5 vision + HQQ (4-bits), weights are in bfloat16
vRAM usage decreased: 8.9 GB -> 3.8 GB
https://colab.research.google.com/drive/176YsmMtdy-o0g19HVUPx7gKZ4OQ2tdxl?usp=sharing
#ai #cv #quantization
vRAM usage decreased: 8.9 GB -> 3.8 GB
https://colab.research.google.com/drive/176YsmMtdy-o0g19HVUPx7gKZ4OQ2tdxl?usp=sharing
#ai #cv #quantization
Google
OCR images using Phi-3.5 vision + HQQ (default backend).ipynb
Colab notebook
doing something
Phi 3.5 vision + HQQ (4-bits), weights are in bfloat16 vRAM usage decreased: 8.9 GB -> 3.8 GB https://colab.research.google.com/drive/176YsmMtdy-o0g19HVUPx7gKZ4OQ2tdxl?usp=sharing #ai #cv #quantization
It is possible to quantize this model as the following:
from transformers import AutoModelForCausalLM, HqqConfig
quant_config = HqqConfig(nbits=4, group_size=64)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="cuda",
torch_dtype=torch.bfloat16,
trust_remote_code=True,
attn_implementation=attn_implementation,
quantization_config=quant_config
)
https://huggingface.co/apple/DepthPro
In this demo I've used Apple's model to extract an object (a man) with transparent background
Also, attached the code you can use to do the same
#ai #cv
In this demo I've used Apple's model to extract an object (a man) with transparent background
Also, attached the code you can use to do the same
#ai #cv
huggingface.co
apple/DepthPro · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.