Hugging Face has a page with available providers and model list they serve
https://huggingface.co/inference/models
It shows the speed in tokens/s so you can choose one for faster inference
https://huggingface.co/inference/models
It shows the speed in tokens/s so you can choose one for faster inference
doing something
Hugging Face has a page with available providers and model list they serve https://huggingface.co/inference/models It shows the speed in tokens/s so you can choose one for faster inference
I've chose LLaMA to check grammar mistakes in my codebase
Time is 1.5 seconds for everything
Time is 1.5 seconds for everything
❤1
Forwarded from Yehor learning Rust
https://github.com/RustedBytes/invoke-llm
Moved away from Python inference script to this one,
Moved away from Python inference script to this one,
cargo install --path . and go to vibe-codeGitHub
GitHub - RustedBytes/invoke-llm: A CLI tool for querying OpenAI-compatible endpoints with a prompt and input file
A CLI tool for querying OpenAI-compatible endpoints with a prompt and input file - RustedBytes/invoke-llm
Yehor learning Rust
https://github.com/RustedBytes/invoke-llm Moved away from Python inference script to this one, cargo install --path . and go to vibe-code
Think about this CLI as a multi-Cursor and sorta dollar-poor way to vibe-code
You can get up a local LLM with
You can get up a local LLM with
llama.cpp and invoke it with your code