doing something
713 subscribers
912 photos
36 videos
36 files
1.94K links
@smlkw doing something, just my notes to keep updates
Download Telegram
Forwarded from Ihor Stepanov
You can reduce the translation step right now thanks to this: https://huggingface.co/knowledgator/gliner-x-large-v0.5
doing something
Photo
Slightly optimized the colab code, now it uses bfloat16 precision

It's more production-ready type of code
🔥1
knowledgator/gliner-bi-base-v1.0 + Helsinki-NLP/opus-mt-tc-big-zle-en in float16 precision

All test cases passing, but some entity types incorrectly detected
If you run MamayLM using llama-cpp:


podman run -v ./models:/models --device="nvidia.com/gpu=0" -p 8000:8000 ghcr.io/ggml-org/llama.cpp:server-cuda -m /models/MamayLM-Gemma-2-9B-IT-v0.1.Q4_K_S.gguf --port 8000 --host 0.0.0.0 -n 1024 --n-gpu-layers -1


Then it uses around 1480MiB on GPU

#nlp #llm
vLLM supports GGUF too - https://docs.vllm.ai/en/v0.9.0.1/features/quantization/gguf.html - but many model architectures are not supported at inference yet

#llm #ai #vllm