Ihor Stepanov
You can reduce the translation step right now thanks to this: https://huggingface.co/knowledgator/gliner-x-large-v0.5
Only one sample was not processed correctly with knowledgator/gliner-x-large-v0.5
But the model is more efficient than previous solution with MamayLM + GLiNER
Colab: https://colab.research.google.com/drive/1tP2gNfvgzscaGKQNkQ0AqBQGSKW6ZUis?usp=sharing
#ai #nlp #pii #security
But the model is more efficient than previous solution with MamayLM + GLiNER
Colab: https://colab.research.google.com/drive/1tP2gNfvgzscaGKQNkQ0AqBQGSKW6ZUis?usp=sharing
#ai #nlp #pii #security
👍2
doing something
Only one sample was not processed correctly with knowledgator/gliner-x-large-v0.5 But the model is more efficient than previous solution with MamayLM + GLiNER Colab: https://colab.research.google.com/drive/1tP2gNfvgzscaGKQNkQ0AqBQGSKW6ZUis?usp=sharing #ai…
Also, a test with the translation model. This hybrid solution should give better results.
Colab: https://colab.research.google.com/drive/1fiOSthyg2BI08x4EKNgZOdGC4cD9DQKe?usp=sharing
#ai #pii #security #nlp
Colab: https://colab.research.google.com/drive/1fiOSthyg2BI08x4EKNgZOdGC4cD9DQKe?usp=sharing
#ai #pii #security #nlp
doing something
Photo
Slightly optimized the colab code, now it uses bfloat16 precision
It's more production-ready type of code
It's more production-ready type of code
🔥1
doing something
Slightly optimized the colab code, now it uses bfloat16 precision It's more production-ready type of code
And we also can optimize the NER model, just use smaller one
👍3
doing something
If you run MamayLM using llama-cpp: podman run -v ./models:/models --device="nvidia.com/gpu=0" -p 8000:8000 ghcr.io/ggml-org/llama.cpp:server-cuda -m /models/MamayLM-Gemma-2-9B-IT-v0.1.Q4_K_S.gguf --port 8000 --host 0.0.0.0 -n 1024 --n-gpu-layers -1 Then…
Experimenting with Gemma 3 and llama-server. gemma-3-1b-it-q4_0.gguf uses around 1350MiB at inference, speed quite high
#llm
#llm
vLLM supports GGUF too - https://docs.vllm.ai/en/v0.9.0.1/features/quantization/gguf.html - but many model architectures are not supported at inference yet
#llm #ai #vllm
#llm #ai #vllm
Forwarded from Feed-Master
YouTube
Nathaniel Simard - Rust for accelerated computing
Recording of a talk given at the Scientific Computing in Rust 2025 online workshop.
This talk highlights how accelerated computing powers AI and explores the design of Burn and CubeCL, leveraging Rust's type system and ownership rules to create flexible…
This talk highlights how accelerated computing powers AI and explores the design of Burn and CubeCL, leveraging Rust's type system and ownership rules to create flexible…
🤯1
doing something
Nice channel, btw #rust
YouTube
Mossa Merhi Reimert - extendr: frictionless bindings for R and Rust
Recording of a talk given at the Scientific Computing in Rust 2024 online workshop.
ExtendR is a suite of software packages that brings Rust to the R ecosystem. First, an overview of the various parts of extendR. (1) libR-sys which is a sys-crate for R API…
ExtendR is a suite of software packages that brings Rust to the R ecosystem. First, an overview of the various parts of extendR. (1) libR-sys which is a sys-crate for R API…
Interesting idea of "code-boarding" using LLMs which explain the code base for new developers
Demo: https://github.com/resemble-ai/chatterbox/blob/4164aa66c46d4a88f06f6cc65a732dcf1ab82db4/.codeboarding/on_boarding.md
#ai #nlp #llm
Demo: https://github.com/resemble-ai/chatterbox/blob/4164aa66c46d4a88f06f6cc65a732dcf1ab82db4/.codeboarding/on_boarding.md
#ai #nlp #llm
GitHub
chatterbox/.codeboarding/on_boarding.md at 4164aa66c46d4a88f06f6cc65a732dcf1ab82db4 · resemble-ai/chatterbox
SoTA open-source TTS. Contribute to resemble-ai/chatterbox development by creating an account on GitHub.
❤1