This media is not supported in your browser
VIEW IN TELEGRAM
📦 edaywalid/undo
Undo
Bring the safety net of ctrl-z directly to your command line. We have all accidentally run a destructive delete or overwritten a file we desperately needed, usually with no way back. This clever tool hooks into your shell to watch destructive system actions in the background. When you run a command, it silently creates instant hardlinks to protect your files before they are modified. If you make a mistake, you just type undo on a new line, and it seamlessly rolls back those exact changes without touching the rest of your system. It is the perfect safety net for every developer.
📰 https://news.ycombinator.com/item?id=49068259
🆔 @hackernewsgithubprojects
Undo
Bring the safety net of ctrl-z directly to your command line. We have all accidentally run a destructive delete or overwritten a file we desperately needed, usually with no way back. This clever tool hooks into your shell to watch destructive system actions in the background. When you run a command, it silently creates instant hardlinks to protect your files before they are modified. If you make a mistake, you just type undo on a new line, and it seamlessly rolls back those exact changes without touching the rest of your system. It is the perfect safety net for every developer.
📰 https://news.ycombinator.com/item?id=49068259
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 zokuzoku/cat-gatekeeper
Cat Gatekeeper
A clever browser extension called Cat Gatekeeper locks you out of distracting websites by forcing you to watch adorable cat videos. Instead of just blocking your favorite social media feeds with a boring error page, this tool tracks your browsing time and interrupts you with a mandatory cat break screen when you have scrolled for too long. It turns the annoying habit of doomscrolling into a gentle, positive reminder to step away from the screen. By replacing frustration with a quick dose of cute animals, it helps you build healthier digital habits without feeling like a punishment.
📰 https://news.ycombinator.com/item?id=49068462
🆔 @hackernewsgithubprojects
Cat Gatekeeper
A clever browser extension called Cat Gatekeeper locks you out of distracting websites by forcing you to watch adorable cat videos. Instead of just blocking your favorite social media feeds with a boring error page, this tool tracks your browsing time and interrupts you with a mandatory cat break screen when you have scrolled for too long. It turns the annoying habit of doomscrolling into a gentle, positive reminder to step away from the screen. By replacing frustration with a quick dose of cute animals, it helps you build healthier digital habits without feeling like a punishment.
📰 https://news.ycombinator.com/item?id=49068462
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 provencher/codex-skills
Codex Skills
Codex Skills is the helper library that lets artificial intelligence break complex jobs into smaller tasks for specialized digital assistants. Instead of trying to tackle a massive project all at once, this system allows a main AI coordinator to handle simple tasks directly while delegating heavier work to focused sub-agents. It runs parallel scouts with low reasoning effort to gather information, assigns routine work to medium-effort agents, and saves high-effort processing for the toughest problems. By keeping these assistants from getting in each other's way, it helps you get cleaner, more organized results without overwhelming a single system.
🆔 @hackernewsgithubprojects
Codex Skills
Codex Skills is the helper library that lets artificial intelligence break complex jobs into smaller tasks for specialized digital assistants. Instead of trying to tackle a massive project all at once, this system allows a main AI coordinator to handle simple tasks directly while delegating heavier work to focused sub-agents. It runs parallel scouts with low reasoning effort to gather information, assigns routine work to medium-effort agents, and saves high-effort processing for the toughest problems. By keeping these assistants from getting in each other's way, it helps you get cleaner, more organized results without overwhelming a single system.
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 nevamind-ai/memu
memu
Give your coding assistants a permanent memory that actually transfers across different apps, sessions, and devices. This clever tool runs quietly in the background, reading your chat transcripts and automatically turning your past coding triumphs into a neat markdown wiki of reusable skills. The next time you start a new session or switch from Claude Code to Cursor, your assistant pulls in that exact knowledge so you never have to explain the same bug or workflow twice. It is basically a shared brain for all your developer tools. Check it out to build your own personal developer memory database.
🆔 @hackernewsgithubprojects
memu
Give your coding assistants a permanent memory that actually transfers across different apps, sessions, and devices. This clever tool runs quietly in the background, reading your chat transcripts and automatically turning your past coding triumphs into a neat markdown wiki of reusable skills. The next time you start a new session or switch from Claude Code to Cursor, your assistant pulls in that exact knowledge so you never have to explain the same bug or workflow twice. It is basically a shared brain for all your developer tools. Check it out to build your own personal developer memory database.
🆔 @hackernewsgithubprojects
Media is too big
VIEW IN TELEGRAM
📦 dadwritestech/llamaforge
LlamaForge
You can now control over two hundred local AI model parameters through a single web browser dashboard instead of typing long terminal commands by hand. LlamaForge is a clean graphical control panel for llama dot cpp and vLLM that runs right on your computer. It reads GGUF metadata directly from your files, lets you fine-tune settings per model, and guides you through building and updating the engine for your specific hardware. The smartest part is its model discovery tool, which rates HuggingFace downloads as fitting, tight, or requiring CPU offload based on your exact graphics memory. It is the ultimate tool to manage and run local artificial intelligence like a...
🆔 @hackernewsgithubprojects
LlamaForge
You can now control over two hundred local AI model parameters through a single web browser dashboard instead of typing long terminal commands by hand. LlamaForge is a clean graphical control panel for llama dot cpp and vLLM that runs right on your computer. It reads GGUF metadata directly from your files, lets you fine-tune settings per model, and guides you through building and updating the engine for your specific hardware. The smartest part is its model discovery tool, which rates HuggingFace downloads as fitting, tight, or requiring CPU offload based on your exact graphics memory. It is the ultimate tool to manage and run local artificial intelligence like a...
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 i7t5/edmund
Edmund Markdown Editor
Stop fighting heavy, sluggish markdown tools and get back to pure writing. Edmund lets you open any markdown file instantly from anywhere on your Mac with zero launch lag, working directly with your local files without forcing you into dedicated vaults or folders. Built natively in Swift using AppKit and TextKit two, this ultra fast editor gives you a beautiful live preview layout while keeping your raw text completely intact. The most impressive part is its custom layout engine, which handles massive multi megabyte files with ease while rendering gorgeous inline math and complex markdown structures perfectly. It is the lightweight, private, and offline first writing companion you have been...
📰 https://news.ycombinator.com/item?id=49070019
🆔 @hackernewsgithubprojects
Edmund Markdown Editor
Stop fighting heavy, sluggish markdown tools and get back to pure writing. Edmund lets you open any markdown file instantly from anywhere on your Mac with zero launch lag, working directly with your local files without forcing you into dedicated vaults or folders. Built natively in Swift using AppKit and TextKit two, this ultra fast editor gives you a beautiful live preview layout while keeping your raw text completely intact. The most impressive part is its custom layout engine, which handles massive multi megabyte files with ease while rendering gorgeous inline math and complex markdown structures perfectly. It is the lightweight, private, and offline first writing companion you have been...
📰 https://news.ycombinator.com/item?id=49070019
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 firish/webfetch
Webfetch
Webfetch is the local search pipeline that gives your AI agents web search capabilities without the massive token bills or subscription fees. While hosted search tools drain your budget by pushing massive pages into your context window, Webfetch runs locally, fetching pages, ranking them, and compressing the text down to only the exact sentences needed. It features an advanced semantic cache that serves paraphrased queries for free, cutting input tokens by up to eighty-eight percent while maintaining hosted accuracy. It is the perfect way to give your local agent loops fast, affordable, and private web access.
🆔 @hackernewsgithubprojects
Webfetch
Webfetch is the local search pipeline that gives your AI agents web search capabilities without the massive token bills or subscription fees. While hosted search tools drain your budget by pushing massive pages into your context window, Webfetch runs locally, fetching pages, ranking them, and compressing the text down to only the exact sentences needed. It features an advanced semantic cache that serves paraphrased queries for free, cutting input tokens by up to eighty-eight percent while maintaining hosted accuracy. It is the perfect way to give your local agent loops fast, affordable, and private web access.
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 nvidia-nemo/labs-oo-agents
NVIDIA Labs Object Oriented Agents
Build and organize AI systems using familiar object-oriented programming rather than juggling messy, disconnected prompts and separate tool frameworks. This project lets you write structured AI agents as standard Python classes where the state lives in object fields, methods define capabilities, and standard type annotations act as strict contracts. A simple ellipsis placeholder in a method body automatically triggers an LLM-driven runtime, meaning the model can write and run Python code in a safe sandbox to complete complex tasks. It makes testing, tracing, and debugging your agents feel exactly like working with traditional, clean software.
📰 https://news.ycombinator.com/item?id=49070316
🆔 @hackernewsgithubprojects
NVIDIA Labs Object Oriented Agents
Build and organize AI systems using familiar object-oriented programming rather than juggling messy, disconnected prompts and separate tool frameworks. This project lets you write structured AI agents as standard Python classes where the state lives in object fields, methods define capabilities, and standard type annotations act as strict contracts. A simple ellipsis placeholder in a method body automatically triggers an LLM-driven runtime, meaning the model can write and run Python code in a safe sandbox to complete complex tasks. It makes testing, tracing, and debugging your agents feel exactly like working with traditional, clean software.
📰 https://news.ycombinator.com/item?id=49070316
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 lightly-ai/lightly-studio
Lightly Studio
lightly-studio is the local browser application that finally unifies your machine learning data workflows into a single cohesive experience. Built with a high-performance Rust backend, it lets you index, curate, and evaluate massive image and video datasets right from your computer. Instead of jumping between disconnected scripts, you can load your files from local folders or cloud storage, instantly visualize ground truth annotations, and compare them directly against model predictions. It also features automatic labeling plugins and visual tools like confusion matrices to help you find where your models fail. It is a incredibly efficient way to clean up your data before training.
🆔 @hackernewsgithubprojects
Lightly Studio
lightly-studio is the local browser application that finally unifies your machine learning data workflows into a single cohesive experience. Built with a high-performance Rust backend, it lets you index, curate, and evaluate massive image and video datasets right from your computer. Instead of jumping between disconnected scripts, you can load your files from local folders or cloud storage, instantly visualize ground truth annotations, and compare them directly against model predictions. It also features automatic labeling plugins and visual tools like confusion matrices to help you find where your models fail. It is a incredibly efficient way to clean up your data before training.
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 microsoft/mlvc
MLVC Neural Video Codec
MLVC is the neural video compressor that runs in real time directly on consumer hardware. While traditional neural video tools are too slow and heavy for everyday devices, this open-source project runs at an average of one hundred frames per second on standard chips from Apple, Intel, and Qualcomm. It slashes video file sizes by more than seventy percent compared to standard hardware encoding, without sacrificing picture quality. The repository comes complete with pre-trained models, training pipelines, and conversion tools to export directly to target devices. It is an amazing leap forward for high-quality streaming on mobile and desktop hardware.
📰 https://news.ycombinator.com/item?id=49070675
🆔 @hackernewsgithubprojects
MLVC Neural Video Codec
MLVC is the neural video compressor that runs in real time directly on consumer hardware. While traditional neural video tools are too slow and heavy for everyday devices, this open-source project runs at an average of one hundred frames per second on standard chips from Apple, Intel, and Qualcomm. It slashes video file sizes by more than seventy percent compared to standard hardware encoding, without sacrificing picture quality. The repository comes complete with pre-trained models, training pipelines, and conversion tools to export directly to target devices. It is an amazing leap forward for high-quality streaming on mobile and desktop hardware.
📰 https://news.ycombinator.com/item?id=49070675
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 gmrandazzo/cheapsecurity
CheapSecurity
Turn a dusty single board computer and a cheap USB webcam into a private home surveillance system. CheapSecurity runs entirely on your own local hardware, so you can keep your video footage safe in your own hands without paying for cloud subscriptions. The system uses smart frame differencing to detect movement, saves clips automatically, and lets you view a live stream from a simple web dashboard. It can even message your phone through Telegram to send you alerts or let you grab a quick snapshot on command. It is a clever, lightweight way to watch over your home on a budget.
📰 https://news.ycombinator.com/item?id=49059398
🆔 @hackernewsgithubprojects
CheapSecurity
Turn a dusty single board computer and a cheap USB webcam into a private home surveillance system. CheapSecurity runs entirely on your own local hardware, so you can keep your video footage safe in your own hands without paying for cloud subscriptions. The system uses smart frame differencing to detect movement, saves clips automatically, and lets you view a live stream from a simple web dashboard. It can even message your phone through Telegram to send you alerts or let you grab a quick snapshot on command. It is a clever, lightweight way to watch over your home on a budget.
📰 https://news.ycombinator.com/item?id=49059398
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 kvcache-ai/agentenv
Agentenv
Spin up thousands of secure, isolated sandboxes for AI agents instantly without melting your servers. Agentenv lets you run massive fleets of Firecracker micro virtual machines that boot or resume in under fifty milliseconds and pause in under one hundred milliseconds. When an agent goes idle, the system automatically reclaims its memory and processing power, and then brings the exact state back online the moment new work arrives. It even supports instant workspace cloning and parallel agent workflows, keeping everything incredibly fast without filling up your local disks. It is the ultimate scaling shortcut for anyone building serious AI agent applications.
🆔 @hackernewsgithubprojects
Agentenv
Spin up thousands of secure, isolated sandboxes for AI agents instantly without melting your servers. Agentenv lets you run massive fleets of Firecracker micro virtual machines that boot or resume in under fifty milliseconds and pause in under one hundred milliseconds. When an agent goes idle, the system automatically reclaims its memory and processing power, and then brings the exact state back online the moment new work arrives. It even supports instant workspace cloning and parallel agent workflows, keeping everything incredibly fast without filling up your local disks. It is the ultimate scaling shortcut for anyone building serious AI agent applications.
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 ronglecat/grok-app
grok-app
The power of the Grok Build command line tool no longer has to be confined to a boring terminal window. A clever unofficial developer tool called grok-app wraps the official Grok agent inside a gorgeous desktop interface. Instead of typing command after command, you get an organized workspace complete with multi-project sessions, visual file previews, and scheduled task automations. It lets you run multiple chats simultaneously, edit resources directly, and manage all your project folders with simple visual controls. It is like giving your local AI developer agent a full-featured control center so you can focus on building instead of managing terminal windows.
🆔 @hackernewsgithubprojects
grok-app
The power of the Grok Build command line tool no longer has to be confined to a boring terminal window. A clever unofficial developer tool called grok-app wraps the official Grok agent inside a gorgeous desktop interface. Instead of typing command after command, you get an organized workspace complete with multi-project sessions, visual file previews, and scheduled task automations. It lets you run multiple chats simultaneously, edit resources directly, and manage all your project folders with simple visual controls. It is like giving your local AI developer agent a full-featured control center so you can focus on building instead of managing terminal windows.
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 alexisfox7/pro-long
Pro-Long AI Memory
Pro-long is the memory expansion for artificial intelligence agents that successfully solves long-term planning tasks without burning through your budget. Traditional setups rely on expensive sub-agents or complex search systems to help AI remember its past actions. This project changes the game by using a simple thirty-line prompt to write every action, observation, and result directly to a single text log. The agent then searches this log using basic code like grep and Python. By treating memory as a simple searchable file, it achieves an incredible ninety-seven point four percent success rate on a top puzzle-solving benchmark at a fraction of the usual cost.
🆔 @hackernewsgithubprojects
Pro-Long AI Memory
Pro-long is the memory expansion for artificial intelligence agents that successfully solves long-term planning tasks without burning through your budget. Traditional setups rely on expensive sub-agents or complex search systems to help AI remember its past actions. This project changes the game by using a simple thirty-line prompt to write every action, observation, and result directly to a single text log. The agent then searches this log using basic code like grep and Python. By treating memory as a simple searchable file, it achieves an incredible ninety-seven point four percent success rate on a top puzzle-solving benchmark at a fraction of the usual cost.
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 patsnap/hiro-moss-ocr
Hiro MOSS OCR
Convert complex document images into clean, structured markup with just a few lines of code. The hiro-moss-ocr project is a compact optical character recognition model that turns tricky visual elements into fully formatted digital text. Instead of giving you messy plain text, it outputs math formulas as math-friendly LaTeX, tables as HTML, and standard paragraphs as clean Markdown. Despite having only about three hundred and twenty million parameters, it handles Japanese, Chinese, and English with ease. It is a lightweight, incredibly efficient tool for anyone looking to extract high-quality, structured data from scans or screenshots without needing massive, power-hungry models.
🆔 @hackernewsgithubprojects
Hiro MOSS OCR
Convert complex document images into clean, structured markup with just a few lines of code. The hiro-moss-ocr project is a compact optical character recognition model that turns tricky visual elements into fully formatted digital text. Instead of giving you messy plain text, it outputs math formulas as math-friendly LaTeX, tables as HTML, and standard paragraphs as clean Markdown. Despite having only about three hundred and twenty million parameters, it handles Japanese, Chinese, and English with ease. It is a lightweight, incredibly efficient tool for anyone looking to extract high-quality, structured data from scans or screenshots without needing massive, power-hungry models.
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 solarkyle/jspace
Inside jspace: Can LLMs Detect Their Own Errors?
Stop relying on basic confidence scores to catch artificial intelligence mistakes. A fascinating new project called jspace analyzes the internal signals of large language models to find out if they actually know when they are hallucinating. By looking deep inside the model's residual stream during a massive twenty-five thousand prompt campaign, researchers trained a tiny three-hundred kilobyte classifier that predicts wrong answers better than the model's own output confidence. The coolest part is that this internal signal successfully transfers across different datasets in the same task family. However, the project's preregistered tests also revealed where these detectors fail, showing they are task-specific rather than universal error monitors. Check out the...
🆔 @hackernewsgithubprojects
Inside jspace: Can LLMs Detect Their Own Errors?
Stop relying on basic confidence scores to catch artificial intelligence mistakes. A fascinating new project called jspace analyzes the internal signals of large language models to find out if they actually know when they are hallucinating. By looking deep inside the model's residual stream during a massive twenty-five thousand prompt campaign, researchers trained a tiny three-hundred kilobyte classifier that predicts wrong answers better than the model's own output confidence. The coolest part is that this internal signal successfully transfers across different datasets in the same task family. However, the project's preregistered tests also revealed where these detectors fail, showing they are task-specific rather than universal error monitors. Check out the...
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 sgshuman/tjs-meal-planner
tjs-meal-planner
Organize your weekly grocery run with a smart assistant that structures your entire menu around a single trip to Trader Joe's. Instead of managing a chaotic daily schedule, this clever tool focuses on a streamlined weekly plan of just one breakfast, one lunch, and two dinners. It automatically sorts your custom shopping list to match the physical walking layout of your local store, so you can breeze through the aisles without backtracking. The best part is the offline support, meaning your list works perfectly even when you lose cell service inside the store. It is the ultimate stress-free shopping companion.
🆔 @hackernewsgithubprojects
tjs-meal-planner
Organize your weekly grocery run with a smart assistant that structures your entire menu around a single trip to Trader Joe's. Instead of managing a chaotic daily schedule, this clever tool focuses on a streamlined weekly plan of just one breakfast, one lunch, and two dinners. It automatically sorts your custom shopping list to match the physical walking layout of your local store, so you can breeze through the aisles without backtracking. The best part is the offline support, meaning your list works perfectly even when you lose cell service inside the store. It is the ultimate stress-free shopping companion.
🆔 @hackernewsgithubprojects
Media is too big
VIEW IN TELEGRAM
📦 mosamlife/wpmgr
WPMgr WordPress Fleet Management
You can now manage and secure an entire fleet of WordPress sites from a single dashboard running on your own server, with zero reliance on third-party cloud services. An open-source tool called wpmgr provides complete self-hosted control over your WordPress installations, handling everything from updates and uptime monitoring to security audits. Its standout feature is a high-performance backup system that uses smart streaming technology to handle massive databases and media libraries without overloading low-power servers. Best of all, it encrypts backups on the client side before they ever leave the site, keeping your data completely private. It is the ultimate dashboard for taking back control of your web infrastructure.
🆔 @hackernewsgithubprojects
WPMgr WordPress Fleet Management
You can now manage and secure an entire fleet of WordPress sites from a single dashboard running on your own server, with zero reliance on third-party cloud services. An open-source tool called wpmgr provides complete self-hosted control over your WordPress installations, handling everything from updates and uptime monitoring to security audits. Its standout feature is a high-performance backup system that uses smart streaming technology to handle massive databases and media libraries without overloading low-power servers. Best of all, it encrypts backups on the client side before they ever leave the site, keeping your data completely private. It is the ultimate dashboard for taking back control of your web infrastructure.
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 aaromanov1985/audio-cassette-simulation
Audio Cassette Simulation
Audio Cassette Simulation is the command line tool that accurately recreates the unique sound profiles of vintage cassette tapes. Using the audio processing power of ffmpeg, this repository applies realistic tape hiss, frequency limits, and pitch wobbles to digital audio files or live web streams. You can make modern music sound like it was recorded on a clean nineties TDK tape, a warm seventies Sony cassette, or even a gritty, highly degraded Soviet MK sixty bootleg. It is an incredibly clever way to instantly inject nostalgic analog warmth and retro character directly into your digital audio tracks.
📰 https://news.ycombinator.com/item?id=49061887
🆔 @hackernewsgithubprojects
Audio Cassette Simulation
Audio Cassette Simulation is the command line tool that accurately recreates the unique sound profiles of vintage cassette tapes. Using the audio processing power of ffmpeg, this repository applies realistic tape hiss, frequency limits, and pitch wobbles to digital audio files or live web streams. You can make modern music sound like it was recorded on a clean nineties TDK tape, a warm seventies Sony cassette, or even a gritty, highly degraded Soviet MK sixty bootleg. It is an incredibly clever way to instantly inject nostalgic analog warmth and retro character directly into your digital audio tracks.
📰 https://news.ycombinator.com/item?id=49061887
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 nikolays/pgsimcity
PGSimCity
PGSimCity is the interactive 3D visualization that finally lets you see exactly how PostgreSQL works under the hood. Instead of dry documentation, it turns database internals into an explorable, living metropolis where every building represents a real mechanism. You can watch client connections stream in, see the shared buffers plaza thrash as pages get evicted, and witness how a single forgotten transaction bloats your storage in real time. It is built for curious developers who want to understand database latency, checkpoints, and replication without drowning in academic jargon. Take a walk through your data directory and master your database today.
📰 https://news.ycombinator.com/item?id=49061947
🆔 @hackernewsgithubprojects
PGSimCity
PGSimCity is the interactive 3D visualization that finally lets you see exactly how PostgreSQL works under the hood. Instead of dry documentation, it turns database internals into an explorable, living metropolis where every building represents a real mechanism. You can watch client connections stream in, see the shared buffers plaza thrash as pages get evicted, and witness how a single forgotten transaction bloats your storage in real time. It is built for curious developers who want to understand database latency, checkpoints, and replication without drowning in academic jargon. Take a walk through your data directory and master your database today.
📰 https://news.ycombinator.com/item?id=49061947
🆔 @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
📦 helasaoudi/llm-inspector
llm-inspector
llm-inspector is the terminal utility that finally explains exactly why your GPU memory is full during local artificial intelligence inference. Instead of just showing a flat VRAM number like standard monitoring tools, this clever utility operates like htop specifically for language models. It probes your active processes to break down your GPU memory into clear categories like weights, key-value cache, and workspace. It even projects future savings from quantization strategies like AWQ or FP8 before you modify your setup. You get a clear, measured breakdown of your hardware limits, helping you pinpoint bottlenecks instantly. Try running the inspect command on your active Ollama or vLLM process to finally see where...
📰 https://news.ycombinator.com/item?id=48956776
🆔 @hackernewsgithubprojects
llm-inspector
llm-inspector is the terminal utility that finally explains exactly why your GPU memory is full during local artificial intelligence inference. Instead of just showing a flat VRAM number like standard monitoring tools, this clever utility operates like htop specifically for language models. It probes your active processes to break down your GPU memory into clear categories like weights, key-value cache, and workspace. It even projects future savings from quantization strategies like AWQ or FP8 before you modify your setup. You get a clear, measured breakdown of your hardware limits, helping you pinpoint bottlenecks instantly. Try running the inspect command on your active Ollama or vLLM process to finally see where...
📰 https://news.ycombinator.com/item?id=48956776
🆔 @hackernewsgithubprojects