https://github.com/EricLBuehler/mistral.rs/blob/master/docs/GEMMA3N.md
Inference engine written in Rust
#rust #gemma
Inference engine written in Rust
#rust #gemma
GitHub
mistral.rs/docs/GEMMA3N.md at master ยท EricLBuehler/mistral.rs
Blazingly fast LLM inference. Contribute to EricLBuehler/mistral.rs development by creating an account on GitHub.
doing something
speed example
On larger machines speed is better (it uses 2 cores), ofc
But for big datasets Rensa requires a lot of memory
16 million of text samples requires ~100 gb of memory
But for big datasets Rensa requires a lot of memory
16 million of text samples requires ~100 gb of memory
HF dataset viewer support chat conversations
Example: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
#hf
Example: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
#hf
My time to use dask, because 16 GB jsonl with non-trivial structure is not fittable into 50 gb of memory...
#data_engineering
#data_engineering
doing something
My time to use dask, because 16 GB jsonl with non-trivial structure is not fittable into 50 gb of memory... #data_engineering
Ohhh, it does not save original structure... List became "Large String"...
Splitted the original file into smaller lines using
Splitted the original file into smaller lines using
split and then processed using polars