HF dataset viewer support chat conversations
Example: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
#hf
Example: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
#hf
My time to use dask, because 16 GB jsonl with non-trivial structure is not fittable into 50 gb of memory...
#data_engineering
#data_engineering
doing something
My time to use dask, because 16 GB jsonl with non-trivial structure is not fittable into 50 gb of memory... #data_engineering
Ohhh, it does not save original structure... List became "Large String"...
Splitted the original file into smaller lines using
Splitted the original file into smaller lines using
split and then processed using polars
doing something
Why good training matters
translator-v02-testlog.txt
126.6 KB
Full log of those translations
https://huggingface.co/spaces/Yehor/en-uk-translator
Now you can play with the model
I will update the checkpoint sometimes
#nlp #llm
Now you can play with the model
I will update the checkpoint sometimes
#nlp #llm
👍2
Forwarded from Задуха
Я тут дізнався, що можу по секрету розповісти про новий НАЙБІЛЬШИЙ корпус української мови з існуючих!
На 60 BILLION tokens!
КОБЗА - https://huggingface.co/datasets/Goader/kobza
Це не просто корпус - це надія, родючий чорнозем, з якого дадуть паростки численні українські ШІ продукти.
Людина яка подарувала нам цей скарб - Mykola Haltiuk (@xgoader), ще і наш недавній підписник!
На 60 BILLION tokens!
КОБЗА - https://huggingface.co/datasets/Goader/kobza
Це не просто корпус - це надія, родючий чорнозем, з якого дадуть паростки численні українські ШІ продукти.
Людина яка подарувала нам цей скарб - Mykola Haltiuk (@xgoader), ще і наш недавній підписник!
huggingface.co
Goader/kobza · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🔥5❤1