doing something
713 subscribers
912 photos
36 videos
36 files
1.94K links
@smlkw doing something, just my notes to keep updates
Download Telegram
If you run MamayLM using llama-cpp:


podman run -v ./models:/models --device="nvidia.com/gpu=0" -p 8000:8000 ghcr.io/ggml-org/llama.cpp:server-cuda -m /models/MamayLM-Gemma-2-9B-IT-v0.1.Q4_K_S.gguf --port 8000 --host 0.0.0.0 -n 1024 --n-gpu-layers -1


Then it uses around 1480MiB on GPU

#nlp #llm
vLLM supports GGUF too - https://docs.vllm.ai/en/v0.9.0.1/features/quantization/gguf.html - but many model architectures are not supported at inference yet

#llm #ai #vllm
GitHub added hashes to binary files

Really useful thing
doing something
https://github.com/mozilla-ocho/llamafile/ #ai
Demo with gemma-3-1b-it-Q6_K model.

Generation is faster, pre-filling is enabled

#llm #ai
🔥1
I think I've seen it before but let's write here as well about this model - https://huggingface.co/osmosis-ai/Osmosis-Structure-0.6B

This model gives an ability to extract information in text using JSON schema

Quantized model (8-bit) requires about 1602MiB of GPU

#ai #llm
👍2🤯1