Forwarded from Yehor learning Rust
https://github.com/RustedBytes/invoke-llm
Moved away from Python inference script to this one,
Moved away from Python inference script to this one,
cargo install --path . and go to vibe-codeGitHub
GitHub - RustedBytes/invoke-llm: A CLI tool for querying OpenAI-compatible endpoints with a prompt and input file
A CLI tool for querying OpenAI-compatible endpoints with a prompt and input file - RustedBytes/invoke-llm
Yehor learning Rust
https://github.com/RustedBytes/invoke-llm Moved away from Python inference script to this one, cargo install --path . and go to vibe-code
Think about this CLI as a multi-Cursor and sorta dollar-poor way to vibe-code
You can get up a local LLM with
You can get up a local LLM with
llama.cpp and invoke it with your codehttps://developer.nvidia.com/blog/delivering-1-5-m-tps-inference-on-nvidia-gb200-nvl72-nvidia-accelerates-openai-gpt-oss-models-from-cloud-to-edge/?ncid=so-twit-199896&linkId=100000376697525
#llm
#llm
NVIDIA Technical Blog
Delivering 1.5 M TPS Inference on NVIDIA GB200 NVL72, NVIDIA Accelerates OpenAI gpt-oss Models from Cloud to Edge
NVIDIA and OpenAI began pushing the boundaries of AI with the launch of NVIDIA DGX back in 2016. The collaborative AI innovation continues with the OpenAI gpt-oss-20b and gpt-oss-120b launch.