SkipCI
333 subscribers
136 photos
169 links
Daily reviews of trending GitHub repos & AI dev tools. Tested, not hyped
Download Telegram
A 744B-parameter model, running on hardware you already own

Frontier MoE models usually need a hyperscaler's GPU cluster. Colibrì streams experts straight from disk into RAM and VRAM, in pure C with zero dependencies, treating storage as part of one memory hierarchy. It boots a 744B-parameter model in about 32 seconds using under 10 GB of resident RAM.

Eight model families run today, from 7B to 2.8T parameters, through one front end: coli chat, coli serve, coli web. A live dashboard shows thousands of experts firing in real time, with token metrics and the VRAM/RAM/disk tier bar.

It's built as an open research platform: no promise on speed, but a hard guarantee that limited fast memory only changes performance, never model semantics. Clone it, point it at a model, and run ./coli chat.

https://github.com/JustVugg/colibri
Stop guessing which LLM your machine can actually run

Downloading a 30GB model only to watch it crawl at 0.5 tok/s is a special kind of pain. llmfit is a Rust CLI that scans your CPU, RAM, GPU, and VRAM, then scores hundreds of open-source models across quality, speed, fit, and context length — before you ever hit download.

It ships as a zero-dependency interactive TUI by default, ranking every model your hardware can handle, plus a CLI and REST API for scripting into pipelines. It works with the runtimes you already use: Ollama, llama.cpp, MLX, LM Studio, and Docker Model Runner. Multi-GPU rigs and MoE architectures are handled too.

The newest trick: benchmark your own runs and submit real tok/s numbers back to the project straight from the TUI. Your hardware's measured results replace estimates for everyone else on the same setup.

Install with brew install AlexsJones/llmfit/llmfit, scoop install llmfit, or the curl installer, then just run llmfit.

https://github.com/AlexsJones/llmfit
An open-source trading agent that never sleeps through a market move

CloddsBot is a self-hosted AI agent that trades across 1000+ markets at once — Polymarket, Kalshi, Binance, Hyperliquid, Solana DEXs, and five EVM chains. Most bots watch one venue and miss everything else; this one scans them all in parallel, finds the edge, and executes without waiting on you.

It's built on Claude, so setup is a conversation rather than a config file: run one command, answer the onboarding wizard, and it's live. Under the hood it manages its own risk — sizing, limits, a kill switch — and ships an agent commerce protocol so it can pay other agents directly, machine to machine.

Written in TypeScript, runs on your own machine, and the whole thing is open for you to read, fork, or point at your own strategies.

github.com/alsk1992/CloddsBot
Most AI research tools skim an abstract and call it done. This one reads the paper.

Hyperresearch turns Claude Code into a research agent that works through dozens to hundreds of sources before writing one sentence, instead of summarizing whatever the search snippet says. Even paywalled papers get chased down through legal open-access routes instead of being cited from a thin auto-generated abstract.

What's unusual is where the work goes: every source it touches lands in a persistent, searchable knowledge base instead of a throwaway context window. Ask about the same topic again next month and it starts from what it already knows rather than researching from zero. The report itself goes through a staged pipeline with adversarial critics attacking the draft before it ships, so claims get checked, not just written.

Install it with pip install hyperresearch, run hyperresearch install inside your project, then type /hyperresearch followed by your topic in Claude Code.

github.com/jordan-gibbs/hyperresearch
Stop making your LLM re-read everything, every time

Most RAG setups retrieve and re-answer from scratch on every single question, burning tokens and never actually accumulating understanding. LLM Wiki reads your documents once and incrementally builds a persistent, interlinked wiki that just keeps growing instead of resetting.

Ingest runs in two steps — the LLM analyzes a source first, then writes wiki pages with full source traceability, and a SHA256 cache skips files that haven't changed. A 4-signal knowledge graph with Louvain community detection surfaces connections you'd never think to search for, and images embedded in your PDFs get captioned and made searchable too.

It's a free cross-platform desktop app with a local HTTP API and bundled MCP server, so it drops straight into Claude Code or Codex as an agent skill and can answer strictly from your own sources.

github.com/nashsu/llm_wiki
Stop running research one experiment at a time

OpenResearch turns Claude Code, Codex, OpenCode, or Cursor into a parallel research team. Instead of one model chasing one idea in one thread, every direction gets its own agent session and its own isolated git worktree, so ten hypotheses can run at once without stepping on each other.

Every run lands in a git-native experiment tree with an immutable, reproducible record, logs and artifacts kept right next to the work that produced them. Point it at local compute, your own infra over SSH, or managed OpenResearch compute, and it can even run the whole loop itself: propose an idea, change the code, launch the experiment, read the evidence, decide what's next. Your code and results stay local by default.

Install the CLI with one command and bring it up locally:

curl -LsSf https://openresearch.sh/install.sh | sh
orx up


That opens a local dashboard where you can watch every parallel run. Written in Rust.

…
Every AI has a hidden prompt. This repo collects them, verbatim.

Behind every polished reply from Claude, ChatGPT, Gemini or Grok sits a long system prompt you never see — the rules, tone, and tool instructions set before your first message. This repo captures those prompts as-is, no paraphrasing, straight from Anthropic, OpenAI, Google, xAI, and more.

The list is wide and current: Claude Fable 5.1, Opus 5 and Claude Code, ChatGPT's GPT-6-Astra and Codex, Gemini 3.8 Flash and Antigravity, Grok, Cursor, Kimi — organized by vendor and updated regularly as new models and tools ship.

It's a plain collection of Markdown files, so browsing, diffing, or grepping across model versions costs nothing. Clone it, or just read a file straight on GitHub.

…
Every AI has a secret instruction manual. This repo collects them.

Before an AI answers your first message, it already read a system prompt telling it how to behave, what to refuse, and how to sound. Normally you never see that text. This repository just publishes it, verbatim, as it leaks out.

The collection spans every major lab: Claude Fable 5.1 and Opus 5, ChatGPT's GPT-6-Astra and Codex, Gemini 3.8 Flash and 3.1 Pro plus Antigravity, Grok and Grok Bot, Cursor, Kimi, and more. Everything is organized by vendor and model name, so you can jump straight to the one you use daily.

It's kept current rather than being a one-time dump — new captures get added as models change, so the same folder structure keeps working as a running archive of how these systems are actually instructed.

To try it, just clone the repo and open the folder for your model — every prompt is a plain, readable file.

github.com/asgeirtj/system_prompts_leaks
A music model that lets you read and edit the song before it's rendered

Most AI song generators are a black box: you type a prompt and get audio you can't touch. YuE2 writes a symbolic melody-and-chord plan first, so you can inspect that plan and change it before anything is rendered into sound.

From the same checkpoint it also handles zero-shot covers — feed it a transcribed song and a new style, and it reworks the arrangement around your melody — and agentic editing, where you describe a harmony or lyric change and it re-renders the full recording to match.

It's a Python package with a small staged API (plan → generate_semantic → synthesize → decode), runs on a single NVIDIA GPU with 24GB VRAM, and outputs 48kHz stereo audio. Clone it, install it, and run the included example script to get your first full song.

https://github.com/multimodal-art-projection/YuE
An AI agent team that hacks your systems so you don't have to

PentAGI runs full penetration tests on its own. No operator babysitting a terminal for weeks — a swarm of specialized AI agents plans the attack, splits it into tasks, and works through recon, exploitation, and reporting end to end.

Every action happens inside an isolated Docker sandbox, wielding 20+ real security tools like nmap, metasploit, and sqlmap. A smart memory system keeps successful approaches for next time, and an optional knowledge graph adds deeper context across a long engagement. Results land as a full vulnerability report with exploitation steps, viewable in the web UI or exported to Markdown and PDF.

It's self-hosted and provider-agnostic: plug in OpenAI, Anthropic, Gemini, Bedrock, Ollama, DeepSeek, and more, plus REST and GraphQL APIs for automation. Written in Go, shipped as a Docker Compose stack — clone it, set your keys, and point it at a target.

…
78 red team skills you can drop straight into Claude

Nobody stays sharp across every attack surface at once — SQL injection, ADCS abuse, EDR evasion, and shellcode all demand different muscle memory. claude-red is a curated library of SKILL.md files, each one loading Claude with expert-level methodology for a single offensive security domain.

The unusual part: skills load on demand from the conversation itself. Mention SQL injection and the relevant skill activates; the rest stay dormant, so you never burn context on techniques you're not using right now. Coverage spans web, wireless, cloud, Active Directory, exploit development, and more.

Getting it running is one command:

git clone https://github.com/SnailSploit/claude-red ~/.claude/skills/claude-red

and Claude behaves like a context-aware operator for whichever attack surface the conversation turns to.

…
Turn a math modeling contest into a one-command paper

MathModelAgent is a Python agent built for math modeling competitions: it analyzes the problem, builds the model, writes and fixes code, and drafts the paper, all in one run. Separate agents handle modeling, coding, and writing, so each stage gets the right model for the job instead of one LLM doing everything.

The output isn't a pile of notes — it's a formatted paper matched to one of 17 contest templates (national and international, Typst-based), backed by a small modeling knowledge base and nine automated checks that catch inconsistent numbers before you submit. Code runs through a local Jupyter interpreter or cloud sandboxes like E2B and daytona, and works with any LLM provider via litellm.

Try the desktop build for a zero-setup start, or run it via Docker, a local Python/Node/Redis install, or as a Claude Code / Codex skill with a single slash command.

https://github.com/jihe520/MathModelAgent
A code review agent that beats Claude Code on precision, at a ninth of the tokens

General-purpose coding agents skimp on large diffs: they skip files, drift on line numbers, and swing wildly with every prompt tweak. Open Code Review fixes that by splitting the job — deterministic engineering decides which files matter and bundles related ones into isolated review units, while an LLM agent does the actual bug hunting with full codebase context.

It comes with a built-in multi-language ruleset for NPE, thread-safety, XSS and SQL injection, and leaves comments precise to the line. It's the same tool that has reviewed code inside Alibaba for two years, across tens of thousands of developers, now open-sourced. Works with any OpenAI- or Anthropic-compatible model endpoint.

Point it at a repo, configure your model of choice, and run ocr on a diff or ocr scan on a whole directory — no fine-tuning, no prompt engineering required.

github.com/alibaba/open-code-review
Your coding assistant just became a video production studio

OpenMontage is the first open-source, agentic video production system: 12 production pipelines, 100+ tools, and over 700 skill and production-knowledge files that teach an AI agent real filmmaking craft. Describe the video in plain language, and the agent handles research, scripting, asset generation, editing, and final composition.

The unusual part: it doesn't just animate a few stills and call it a video. The agent can build a corpus from free stock footage and open archives, retrieve real motion clips, edit them into a timeline, and render a finished cut — an actual production pipeline, not a slideshow trick.

It's written in Python and runs on top of whatever AI coding assistant you already use. Clone the repo, point your agent at it, and start from a prompt or from a video you already love.

github.com/calesthio/OpenMontage
See through walls using nothing but WiFi

Cameras feel invasive, wearables get lost or forgotten. RuView reads the radio reflections your WiFi router is already producing and turns them into presence detection, movement tracking, and contactless vital signs — breathing and heart rate — without a single pixel of video.

The sensor is a $9 ESP32 board capturing Channel State Information; a pretrained model just 8 KB in size (4-bit quantized) turns that signal into data in microseconds, running on hardware as small as a Raspberry Pi, with no cloud and no internet connection needed. It drops straight into Home Assistant, Apple Home, Google Home, and Alexa.

Written in Rust, MIT-licensed, with over a thousand tests passing. Clone it, flash an ESP32, and give any room spatial awareness.

github.com/ruvnet/RuView
A self-hosted box that holds Wikipedia, thousands of books, and AI — with zero internet

Project NOMAD is an offline-first knowledge and education server. Once installed, it needs no connection at all: full Wikipedia, medical references, ebooks, courses, and regional maps all live on hardware you own.

What's unusual is the scope: it bundles a whole stack — Kiwix for the library, Kolibri for Khan Academy-style courses, ProtoMaps for maps, CyberChef for data tools, and an optional local AI assistant with document upload and semantic search, running through Ollama or any OpenAI-compatible server. One "Command Center" UI manages all of it, including updates, on a schedule you control.

Setup is a single install script on any Debian-based OS, or a Docker Compose file if you want manual control. No desktop environment required — everything, including the AI chat, runs through the browser once installed. Minimum specs are modest; a GPU helps if you want to run larger local models.

…
Turn a spare iPhone into a real second monitor for your Mac — free

OpenDisplay is an open-source alternative to Sidecar, Duet, and Luna: a true extended display, not a mirror, with no subscription, no dongle, and no shared-Apple-ID requirement. Drag a window onto the phone and it just lives there.

It streams over USB or WiFi, matches the device's panel pixel-for-pixel for Retina-sharp text, and rebuilds itself for portrait or landscape. Touch works like a trackpad: tap to click, drag to drag, two-finger scroll — all over a low-latency H.264 pipeline and a single direct TCP connection, no server involved.

Written in Swift, GPL-3.0, and built to be read and extended: the wire protocol is documented, so anyone can write a new client instead of reverse-engineering the app — even an old Mac can become a display.

…
Homebrew finally has an official app — and it never hides the terminal from you

BrewUI is Homebrew's own macOS GUI, built for people who want to search, install, update, and manage packages without memorizing CLI incantations. It's a native SwiftUI app, not a wrapper that pretends the command line doesn't exist.

The whole point is transparency: every action runs the real brew CLI underneath, and a live console shows the exact commands as they execute. Nothing is simulated, nothing is hidden — you get a graphical front end with zero mystery about what's happening to your system.

Configuration stays outside your shell entirely. BrewUI launches Homebrew through a clean, isolated environment and reads settings from dedicated brew.env files instead of your aliases or exported variables, so behavior stays predictable regardless of your shell setup.

Install it straight from Homebrew itself:
brew install --cask homebrew-app

github.com/Homebrew/BrewUI
Paste a link, get the file — Udemy courses, YouTube, 1,800+ sites, no terminal

You bought a course and want it saved before the platform pulls it. You keep a yt-dlp cheat sheet because the flags never stick. You have five different tools for Instagram, X, Pinterest and torrents, and none of them remember your login. OmniGet puts all of that behind one text box: paste a link, pick a quality, download.

It doesn't stop at the download. The same window plays the course you just grabbed, opens the PDF or EPUB, and manages your music library. Under the hood it runs on yt-dlp, which installs and updates itself, so there's nothing to configure. It's a free, open-source desktop app for Windows, macOS and Linux, and your files never leave your computer — no account, no ads, no telemetry.

Grab a build from the releases page and try it on whatever's sitting in your clipboard right now.

github.com/tonhowtf/omniget
The NSA open-sourced its own reverse engineering tool. It's free forever.

Ghidra is a full software reverse engineering framework: disassembler, decompiler, and graph view in one place, so you can take any compiled binary and turn it back into readable code. Analysis work that used to require expensive commercial licenses now runs on a free download.

It handles a wide range of processor architectures and executable formats, works on Windows, macOS, and Linux, and scales from quick interactive lookups to fully automated pipelines. Beyond the built-in tools, you can write your own analysis scripts and extensions in Java or Python.

Grab a release, install a JDK, and run ./ghidraRun to launch it — or build straight from source with Gradle if you want the latest development version.

https://github.com/NationalSecurityAgency/ghidra
A macOS launcher that stays under 100 MB, because it never left Swift

Tinycast is a fully native launcher, hotkey system, and clipboard history for macOS — SwiftUI and AppKit, zero third-party dependencies, no Electron shell, no telemetry. One global hotkey opens the palette from anywhere: fuzzy-search apps, files, and clipboard text and images, run inline currency and crypto conversions, or snap windows into halves and thirds, all in a footprint most launchers spend on their splash screen.

The unusual part: it runs your existing Raycast extensions natively, rendered as real SwiftUI rather than a web view. Snippets, quicklinks, Apple Shortcuts, and custom shell commands round it out, all searchable from the same palette.

Install it with Homebrew — brew tap abue-ammar/tinycast, then brew install --cask tinycast for Apple silicon on macOS 26+. It's free and open source under AGPL-3.0.

https://github.com/abue-ammar/tinycast