Your AI agent can write HTML all day. Now it can ship video too.
HyperFrames is an open-source framework that renders HTML, CSS, media, and seekable animations straight into deterministic MP4 files. No editor, no timeline — just markup in, video out, the same result every time.
It's built for agents, not humans clicking timelines. Skills teach Claude Code, Cursor, Codex, Gemini CLI, and other agents the full loop: plan the video, write valid HTML, wire up animations, add media, lint, preview, render. One prompt describing a video is enough to trigger it.
Try it with
…
HyperFrames is an open-source framework that renders HTML, CSS, media, and seekable animations straight into deterministic MP4 files. No editor, no timeline — just markup in, video out, the same result every time.
It's built for agents, not humans clicking timelines. Skills teach Claude Code, Cursor, Codex, Gemini CLI, and other agents the full loop: plan the video, write valid HTML, wire up animations, add media, lint, preview, render. One prompt describing a video is enough to trigger it.
Try it with
npx skills add heygen-com/hyperframes, or for agent and non-interactive runs npx hyperframes skills update. Written in TypeScript, Apache-2.0 licensed.…
Your AI agent is burning its own context — this MCP server stops it
Every tool call dumps raw data into your agent's context window — a single Playwright snapshot alone costs 56 KB. After half an hour, 40% of that context can be gone, and when the conversation compacts, the agent forgets what it was doing.
Context Mode is an MCP server that sandboxes tool output before it reaches the model — 315 KB drops to 5.4 KB, a 98% cut — and persists session memory in SQLite with FTS5 search, so a resumed session picks up exactly where it left off.
It also changes the habit: instead of reading dozens of files into context, the agent writes a small script that does the work and logs only the result. Routing is enforced via hooks across 17 platforms, not left as something to remember.
Install through Claude Code's plugin marketplace or npm, then run the built-in doctor command to confirm everything is wired up.
mksglu/context-mode
Every tool call dumps raw data into your agent's context window — a single Playwright snapshot alone costs 56 KB. After half an hour, 40% of that context can be gone, and when the conversation compacts, the agent forgets what it was doing.
Context Mode is an MCP server that sandboxes tool output before it reaches the model — 315 KB drops to 5.4 KB, a 98% cut — and persists session memory in SQLite with FTS5 search, so a resumed session picks up exactly where it left off.
It also changes the habit: instead of reading dozens of files into context, the agent writes a small script that does the work and logs only the result. Routing is enforced via hooks across 17 platforms, not left as something to remember.
Install through Claude Code's plugin marketplace or npm, then run the built-in doctor command to confirm everything is wired up.
mksglu/context-mode
Stop feeding your LLM raw PDFs
Messy PDFs, Word docs, and spreadsheets are a mess for language models to parse. MarkItDown is a lightweight Python tool that converts them into clean Markdown, keeping the structure that matters: headings, lists, tables, and links, instead of a wall of unformatted text.
It goes further than office files. Images get OCR and EXIF metadata, audio gets transcribed, and YouTube URLs get their captions pulled — all normalized into the same Markdown your model already reads natively. ZIP archives are handled too, converting everything inside.
Trying it takes one line:
github.com/microsoft/markitdown
Messy PDFs, Word docs, and spreadsheets are a mess for language models to parse. MarkItDown is a lightweight Python tool that converts them into clean Markdown, keeping the structure that matters: headings, lists, tables, and links, instead of a wall of unformatted text.
It goes further than office files. Images get OCR and EXIF metadata, audio gets transcribed, and YouTube URLs get their captions pulled — all normalized into the same Markdown your model already reads natively. ZIP archives are handled too, converting everything inside.
Trying it takes one line:
pip install 'markitdown[all]', then markitdown path-to-file.pdf > document.md. Pipe input in, get Markdown out. A plugin system adds extras like LLM-vision OCR for scanned documents.github.com/microsoft/markitdown
Give your AI agent hands: it clicks, types, and fills forms like a human
Most agents choke the moment a task needs a real browser: pop-ups, logins, multi-step forms, messy layouts. Browser Use hands your agent an actual browser instead of a text box, so it can open pages, click buttons, and complete the task you described in plain English.
It works both ways: point it at a job application and it fills every field and submits, or point it at a profile page and it extracts structured data straight to CSV. Pick any LLM, run it locally, and pair it with any agent you already use, from Claude Code to Cursor.
Install with
https://github.com/browser-use/browser-use
Most agents choke the moment a task needs a real browser: pop-ups, logins, multi-step forms, messy layouts. Browser Use hands your agent an actual browser instead of a text box, so it can open pages, click buttons, and complete the task you described in plain English.
It works both ways: point it at a job application and it fills every field and submits, or point it at a profile page and it extracts structured data straight to CSV. Pick any LLM, run it locally, and pair it with any agent you already use, from Claude Code to Cursor.
Install with
uv add browser-use or pip install browser-use, drop an API key in .env, and run your first agent in a few lines of Python. It's open source and written in Python, ready to plug into whatever you're building.https://github.com/browser-use/browser-use
An AI trading desk you can run from a single Python repo
TradingAgents is a multi-agent framework that simulates a trading firm instead of one model doing everything. Fundamentals, sentiment, news, and technical analysts each produce their own read, then bullish and bearish researchers debate the findings before a trader and a risk-management team decide whether the trade happens.
The interesting part is the disagreement: agents argue through structured debate rounds, and a portfolio manager approves or rejects the final proposal before anything reaches the (simulated) exchange. It supports many LLM providers, from GPT and Claude to Gemini, Grok, DeepSeek, Qwen, and local Ollama models.
Clone it, add your API keys, and run the CLI or import it as a package. Built for research, not financial advice, but a way to see how far agent teams get on messy, real-world data.
https://github.com/TauricResearch/TradingAgents
TradingAgents is a multi-agent framework that simulates a trading firm instead of one model doing everything. Fundamentals, sentiment, news, and technical analysts each produce their own read, then bullish and bearish researchers debate the findings before a trader and a risk-management team decide whether the trade happens.
The interesting part is the disagreement: agents argue through structured debate rounds, and a portfolio manager approves or rejects the final proposal before anything reaches the (simulated) exchange. It supports many LLM providers, from GPT and Claude to Gemini, Grok, DeepSeek, Qwen, and local Ollama models.
Clone it, add your API keys, and run the CLI or import it as a package. Built for research, not financial advice, but a way to see how far agent teams get on messy, real-world data.
https://github.com/TauricResearch/TradingAgents
Design 3D buildings in your browser, no CAD required
Pascal Editor is an open-source 3D building editor that runs entirely in the browser, built with React Three Fiber and WebGPU. Sketch walls, drop in furniture, and render the result instantly, without months of learning professional CAD software.
What sets it apart: AI agents can build scenes too. Pascal ships an MCP server, so an agent can connect and design a scene directly, with dedicated agent skills for scene setup and furniture-fit checks.
The whole thing is a TypeScript monorepo under the MIT license, split into clean packages, a core state layer, a rendering viewer, the editing tools, and a CLI that ties it all together.
Try it locally with a single command:
It starts the editor plus an authenticated MCP service on your machine, no repository clone needed.
https://github.com/pascalorg/editor
Pascal Editor is an open-source 3D building editor that runs entirely in the browser, built with React Three Fiber and WebGPU. Sketch walls, drop in furniture, and render the result instantly, without months of learning professional CAD software.
What sets it apart: AI agents can build scenes too. Pascal ships an MCP server, so an agent can connect and design a scene directly, with dedicated agent skills for scene setup and furniture-fit checks.
The whole thing is a TypeScript monorepo under the MIT license, split into clean packages, a core state layer, a rendering viewer, the editing tools, and a CLI that ties it all together.
Try it locally with a single command:
npx @pascal-app/cli editorIt starts the editor plus an authenticated MCP service on your machine, no repository clone needed.
https://github.com/pascalorg/editor
❤1
Skills that let your coding agent design real hardware
Most coding agents can write software but stall the moment a task needs actual geometry. text-to-cad is a Python library of agent skills that plugs that gap: it generates and edits CAD models from plain-language requests, exporting STEP, STL, 3MF, or GLB.
Beyond CAD, the library covers the rest of the pipeline: sourcing off-the-shelf STEP parts, drawing 2D DXF layouts, writing URDF/SRDF/SDF robot description files, checking mesh printability, and slicing meshes into printer-ready G-code.
Install it with the Skills CLI (
…
Most coding agents can write software but stall the moment a task needs actual geometry. text-to-cad is a Python library of agent skills that plugs that gap: it generates and edits CAD models from plain-language requests, exporting STEP, STL, 3MF, or GLB.
Beyond CAD, the library covers the rest of the pipeline: sourcing off-the-shelf STEP parts, drawing 2D DXF layouts, writing URDF/SRDF/SDF robot description files, checking mesh printability, and slicing meshes into printer-ready G-code.
Install it with the Skills CLI (
npx skills add earthtojake/text-to-cad) for supported agents, or grab the native plugin for Codex, Claude Code, or Grok Build. It's MIT-licensed and open source.…
Your browser just became a spy satellite — and every feed on it is real
God's Eye View puts live aircraft, ships, satellites, earthquakes, and public cameras onto one photorealistic 3D globe, running locally in your browser. No dashboard full of tabs, no separate trackers for planes and ships and quakes — one map, real public data, click anything to follow it.
The unusual part: click a plane and you ride inside its cockpit as the camera tracks the terrain beneath it. Switch the whole globe to thermal, night vision, or a military HUD with a GLSL shader swap. There's also a realtime voice agent, so you can just talk to it and have it draw boundaries or routes on the world for you.
It runs keyless out of the box — clone it,
…
God's Eye View puts live aircraft, ships, satellites, earthquakes, and public cameras onto one photorealistic 3D globe, running locally in your browser. No dashboard full of tabs, no separate trackers for planes and ships and quakes — one map, real public data, click anything to follow it.
The unusual part: click a plane and you ride inside its cockpit as the camera tracks the terrain beneath it. Switch the whole globe to thermal, night vision, or a military HUD with a GLSL shader swap. There's also a realtime voice agent, so you can just talk to it and have it draw boundaries or routes on the world for you.
It runs keyless out of the box — clone it,
npm ci, npm run dev, and you're looking at live traffic with no accounts and no API keys. Add a free Cesium ion token later for full photorealistic terrain if you want it. It's plain JavaScript, fully open source, and built to be extended with your own data layers.…
A 744B-parameter model, running on hardware you already own
Frontier MoE models usually need a hyperscaler's GPU cluster. Colibrì streams experts straight from disk into RAM and VRAM, in pure C with zero dependencies, treating storage as part of one memory hierarchy. It boots a 744B-parameter model in about 32 seconds using under 10 GB of resident RAM.
Eight model families run today, from 7B to 2.8T parameters, through one front end:
It's built as an open research platform: no promise on speed, but a hard guarantee that limited fast memory only changes performance, never model semantics. Clone it, point it at a model, and run
https://github.com/JustVugg/colibri
Frontier MoE models usually need a hyperscaler's GPU cluster. Colibrì streams experts straight from disk into RAM and VRAM, in pure C with zero dependencies, treating storage as part of one memory hierarchy. It boots a 744B-parameter model in about 32 seconds using under 10 GB of resident RAM.
Eight model families run today, from 7B to 2.8T parameters, through one front end:
coli chat, coli serve, coli web. A live dashboard shows thousands of experts firing in real time, with token metrics and the VRAM/RAM/disk tier bar.It's built as an open research platform: no promise on speed, but a hard guarantee that limited fast memory only changes performance, never model semantics. Clone it, point it at a model, and run
./coli chat.https://github.com/JustVugg/colibri
Stop guessing which LLM your machine can actually run
Downloading a 30GB model only to watch it crawl at 0.5 tok/s is a special kind of pain.
It ships as a zero-dependency interactive TUI by default, ranking every model your hardware can handle, plus a CLI and REST API for scripting into pipelines. It works with the runtimes you already use: Ollama, llama.cpp, MLX, LM Studio, and Docker Model Runner. Multi-GPU rigs and MoE architectures are handled too.
The newest trick: benchmark your own runs and submit real tok/s numbers back to the project straight from the TUI. Your hardware's measured results replace estimates for everyone else on the same setup.
Install with
https://github.com/AlexsJones/llmfit
Downloading a 30GB model only to watch it crawl at 0.5 tok/s is a special kind of pain.
llmfit is a Rust CLI that scans your CPU, RAM, GPU, and VRAM, then scores hundreds of open-source models across quality, speed, fit, and context length — before you ever hit download.It ships as a zero-dependency interactive TUI by default, ranking every model your hardware can handle, plus a CLI and REST API for scripting into pipelines. It works with the runtimes you already use: Ollama, llama.cpp, MLX, LM Studio, and Docker Model Runner. Multi-GPU rigs and MoE architectures are handled too.
The newest trick: benchmark your own runs and submit real tok/s numbers back to the project straight from the TUI. Your hardware's measured results replace estimates for everyone else on the same setup.
Install with
brew install AlexsJones/llmfit/llmfit, scoop install llmfit, or the curl installer, then just run llmfit.https://github.com/AlexsJones/llmfit
An open-source trading agent that never sleeps through a market move
CloddsBot is a self-hosted AI agent that trades across 1000+ markets at once — Polymarket, Kalshi, Binance, Hyperliquid, Solana DEXs, and five EVM chains. Most bots watch one venue and miss everything else; this one scans them all in parallel, finds the edge, and executes without waiting on you.
It's built on Claude, so setup is a conversation rather than a config file: run one command, answer the onboarding wizard, and it's live. Under the hood it manages its own risk — sizing, limits, a kill switch — and ships an agent commerce protocol so it can pay other agents directly, machine to machine.
Written in TypeScript, runs on your own machine, and the whole thing is open for you to read, fork, or point at your own strategies.
github.com/alsk1992/CloddsBot
CloddsBot is a self-hosted AI agent that trades across 1000+ markets at once — Polymarket, Kalshi, Binance, Hyperliquid, Solana DEXs, and five EVM chains. Most bots watch one venue and miss everything else; this one scans them all in parallel, finds the edge, and executes without waiting on you.
It's built on Claude, so setup is a conversation rather than a config file: run one command, answer the onboarding wizard, and it's live. Under the hood it manages its own risk — sizing, limits, a kill switch — and ships an agent commerce protocol so it can pay other agents directly, machine to machine.
Written in TypeScript, runs on your own machine, and the whole thing is open for you to read, fork, or point at your own strategies.
github.com/alsk1992/CloddsBot
Most AI research tools skim an abstract and call it done. This one reads the paper.
Hyperresearch turns Claude Code into a research agent that works through dozens to hundreds of sources before writing one sentence, instead of summarizing whatever the search snippet says. Even paywalled papers get chased down through legal open-access routes instead of being cited from a thin auto-generated abstract.
What's unusual is where the work goes: every source it touches lands in a persistent, searchable knowledge base instead of a throwaway context window. Ask about the same topic again next month and it starts from what it already knows rather than researching from zero. The report itself goes through a staged pipeline with adversarial critics attacking the draft before it ships, so claims get checked, not just written.
Install it with
github.com/jordan-gibbs/hyperresearch
Hyperresearch turns Claude Code into a research agent that works through dozens to hundreds of sources before writing one sentence, instead of summarizing whatever the search snippet says. Even paywalled papers get chased down through legal open-access routes instead of being cited from a thin auto-generated abstract.
What's unusual is where the work goes: every source it touches lands in a persistent, searchable knowledge base instead of a throwaway context window. Ask about the same topic again next month and it starts from what it already knows rather than researching from zero. The report itself goes through a staged pipeline with adversarial critics attacking the draft before it ships, so claims get checked, not just written.
Install it with
pip install hyperresearch, run hyperresearch install inside your project, then type /hyperresearch followed by your topic in Claude Code.github.com/jordan-gibbs/hyperresearch
Stop making your LLM re-read everything, every time
Most RAG setups retrieve and re-answer from scratch on every single question, burning tokens and never actually accumulating understanding. LLM Wiki reads your documents once and incrementally builds a persistent, interlinked wiki that just keeps growing instead of resetting.
Ingest runs in two steps — the LLM analyzes a source first, then writes wiki pages with full source traceability, and a SHA256 cache skips files that haven't changed. A 4-signal knowledge graph with Louvain community detection surfaces connections you'd never think to search for, and images embedded in your PDFs get captioned and made searchable too.
It's a free cross-platform desktop app with a local HTTP API and bundled MCP server, so it drops straight into Claude Code or Codex as an agent skill and can answer strictly from your own sources.
github.com/nashsu/llm_wiki
Most RAG setups retrieve and re-answer from scratch on every single question, burning tokens and never actually accumulating understanding. LLM Wiki reads your documents once and incrementally builds a persistent, interlinked wiki that just keeps growing instead of resetting.
Ingest runs in two steps — the LLM analyzes a source first, then writes wiki pages with full source traceability, and a SHA256 cache skips files that haven't changed. A 4-signal knowledge graph with Louvain community detection surfaces connections you'd never think to search for, and images embedded in your PDFs get captioned and made searchable too.
It's a free cross-platform desktop app with a local HTTP API and bundled MCP server, so it drops straight into Claude Code or Codex as an agent skill and can answer strictly from your own sources.
github.com/nashsu/llm_wiki
Stop running research one experiment at a time
OpenResearch turns Claude Code, Codex, OpenCode, or Cursor into a parallel research team. Instead of one model chasing one idea in one thread, every direction gets its own agent session and its own isolated git worktree, so ten hypotheses can run at once without stepping on each other.
Every run lands in a git-native experiment tree with an immutable, reproducible record, logs and artifacts kept right next to the work that produced them. Point it at local compute, your own infra over SSH, or managed OpenResearch compute, and it can even run the whole loop itself: propose an idea, change the code, launch the experiment, read the evidence, decide what's next. Your code and results stay local by default.
Install the CLI with one command and bring it up locally:
That opens a local dashboard where you can watch every parallel run. Written in Rust.
…
OpenResearch turns Claude Code, Codex, OpenCode, or Cursor into a parallel research team. Instead of one model chasing one idea in one thread, every direction gets its own agent session and its own isolated git worktree, so ten hypotheses can run at once without stepping on each other.
Every run lands in a git-native experiment tree with an immutable, reproducible record, logs and artifacts kept right next to the work that produced them. Point it at local compute, your own infra over SSH, or managed OpenResearch compute, and it can even run the whole loop itself: propose an idea, change the code, launch the experiment, read the evidence, decide what's next. Your code and results stay local by default.
Install the CLI with one command and bring it up locally:
curl -LsSf https://openresearch.sh/install.sh | sh
orx upThat opens a local dashboard where you can watch every parallel run. Written in Rust.
…
Every AI has a hidden prompt. This repo collects them, verbatim.
Behind every polished reply from Claude, ChatGPT, Gemini or Grok sits a long system prompt you never see — the rules, tone, and tool instructions set before your first message. This repo captures those prompts as-is, no paraphrasing, straight from Anthropic, OpenAI, Google, xAI, and more.
The list is wide and current: Claude Fable 5.1, Opus 5 and Claude Code, ChatGPT's GPT-6-Astra and Codex, Gemini 3.8 Flash and Antigravity, Grok, Cursor, Kimi — organized by vendor and updated regularly as new models and tools ship.
It's a plain collection of Markdown files, so browsing, diffing, or grepping across model versions costs nothing. Clone it, or just read a file straight on GitHub.
…
Behind every polished reply from Claude, ChatGPT, Gemini or Grok sits a long system prompt you never see — the rules, tone, and tool instructions set before your first message. This repo captures those prompts as-is, no paraphrasing, straight from Anthropic, OpenAI, Google, xAI, and more.
The list is wide and current: Claude Fable 5.1, Opus 5 and Claude Code, ChatGPT's GPT-6-Astra and Codex, Gemini 3.8 Flash and Antigravity, Grok, Cursor, Kimi — organized by vendor and updated regularly as new models and tools ship.
It's a plain collection of Markdown files, so browsing, diffing, or grepping across model versions costs nothing. Clone it, or just read a file straight on GitHub.
…
Every AI has a secret instruction manual. This repo collects them.
Before an AI answers your first message, it already read a system prompt telling it how to behave, what to refuse, and how to sound. Normally you never see that text. This repository just publishes it, verbatim, as it leaks out.
The collection spans every major lab: Claude Fable 5.1 and Opus 5, ChatGPT's GPT-6-Astra and Codex, Gemini 3.8 Flash and 3.1 Pro plus Antigravity, Grok and Grok Bot, Cursor, Kimi, and more. Everything is organized by vendor and model name, so you can jump straight to the one you use daily.
It's kept current rather than being a one-time dump — new captures get added as models change, so the same folder structure keeps working as a running archive of how these systems are actually instructed.
To try it, just clone the repo and open the folder for your model — every prompt is a plain, readable file.
github.com/asgeirtj/system_prompts_leaks
Before an AI answers your first message, it already read a system prompt telling it how to behave, what to refuse, and how to sound. Normally you never see that text. This repository just publishes it, verbatim, as it leaks out.
The collection spans every major lab: Claude Fable 5.1 and Opus 5, ChatGPT's GPT-6-Astra and Codex, Gemini 3.8 Flash and 3.1 Pro plus Antigravity, Grok and Grok Bot, Cursor, Kimi, and more. Everything is organized by vendor and model name, so you can jump straight to the one you use daily.
It's kept current rather than being a one-time dump — new captures get added as models change, so the same folder structure keeps working as a running archive of how these systems are actually instructed.
To try it, just clone the repo and open the folder for your model — every prompt is a plain, readable file.
github.com/asgeirtj/system_prompts_leaks
A music model that lets you read and edit the song before it's rendered
Most AI song generators are a black box: you type a prompt and get audio you can't touch. YuE2 writes a symbolic melody-and-chord plan first, so you can inspect that plan and change it before anything is rendered into sound.
From the same checkpoint it also handles zero-shot covers — feed it a transcribed song and a new style, and it reworks the arrangement around your melody — and agentic editing, where you describe a harmony or lyric change and it re-renders the full recording to match.
It's a Python package with a small staged API (plan → generate_semantic → synthesize → decode), runs on a single NVIDIA GPU with 24GB VRAM, and outputs 48kHz stereo audio. Clone it, install it, and run the included example script to get your first full song.
https://github.com/multimodal-art-projection/YuE
Most AI song generators are a black box: you type a prompt and get audio you can't touch. YuE2 writes a symbolic melody-and-chord plan first, so you can inspect that plan and change it before anything is rendered into sound.
From the same checkpoint it also handles zero-shot covers — feed it a transcribed song and a new style, and it reworks the arrangement around your melody — and agentic editing, where you describe a harmony or lyric change and it re-renders the full recording to match.
It's a Python package with a small staged API (plan → generate_semantic → synthesize → decode), runs on a single NVIDIA GPU with 24GB VRAM, and outputs 48kHz stereo audio. Clone it, install it, and run the included example script to get your first full song.
https://github.com/multimodal-art-projection/YuE
An AI agent team that hacks your systems so you don't have to
PentAGI runs full penetration tests on its own. No operator babysitting a terminal for weeks — a swarm of specialized AI agents plans the attack, splits it into tasks, and works through recon, exploitation, and reporting end to end.
Every action happens inside an isolated Docker sandbox, wielding 20+ real security tools like
It's self-hosted and provider-agnostic: plug in OpenAI, Anthropic, Gemini, Bedrock, Ollama, DeepSeek, and more, plus REST and GraphQL APIs for automation. Written in Go, shipped as a Docker Compose stack — clone it, set your keys, and point it at a target.
…
PentAGI runs full penetration tests on its own. No operator babysitting a terminal for weeks — a swarm of specialized AI agents plans the attack, splits it into tasks, and works through recon, exploitation, and reporting end to end.
Every action happens inside an isolated Docker sandbox, wielding 20+ real security tools like
nmap, metasploit, and sqlmap. A smart memory system keeps successful approaches for next time, and an optional knowledge graph adds deeper context across a long engagement. Results land as a full vulnerability report with exploitation steps, viewable in the web UI or exported to Markdown and PDF.It's self-hosted and provider-agnostic: plug in OpenAI, Anthropic, Gemini, Bedrock, Ollama, DeepSeek, and more, plus REST and GraphQL APIs for automation. Written in Go, shipped as a Docker Compose stack — clone it, set your keys, and point it at a target.
…
78 red team skills you can drop straight into Claude
Nobody stays sharp across every attack surface at once — SQL injection, ADCS abuse, EDR evasion, and shellcode all demand different muscle memory.
The unusual part: skills load on demand from the conversation itself. Mention SQL injection and the relevant skill activates; the rest stay dormant, so you never burn context on techniques you're not using right now. Coverage spans web, wireless, cloud, Active Directory, exploit development, and more.
Getting it running is one command:
and Claude behaves like a context-aware operator for whichever attack surface the conversation turns to.
…
Nobody stays sharp across every attack surface at once — SQL injection, ADCS abuse, EDR evasion, and shellcode all demand different muscle memory.
claude-red is a curated library of SKILL.md files, each one loading Claude with expert-level methodology for a single offensive security domain.The unusual part: skills load on demand from the conversation itself. Mention SQL injection and the relevant skill activates; the rest stay dormant, so you never burn context on techniques you're not using right now. Coverage spans web, wireless, cloud, Active Directory, exploit development, and more.
Getting it running is one command:
git clone https://github.com/SnailSploit/claude-red ~/.claude/skills/claude-redand Claude behaves like a context-aware operator for whichever attack surface the conversation turns to.
…
Turn a math modeling contest into a one-command paper
MathModelAgent is a Python agent built for math modeling competitions: it analyzes the problem, builds the model, writes and fixes code, and drafts the paper, all in one run. Separate agents handle modeling, coding, and writing, so each stage gets the right model for the job instead of one LLM doing everything.
The output isn't a pile of notes — it's a formatted paper matched to one of 17 contest templates (national and international, Typst-based), backed by a small modeling knowledge base and nine automated checks that catch inconsistent numbers before you submit. Code runs through a local Jupyter interpreter or cloud sandboxes like E2B and daytona, and works with any LLM provider via litellm.
Try the desktop build for a zero-setup start, or run it via Docker, a local Python/Node/Redis install, or as a Claude Code / Codex skill with a single slash command.
https://github.com/jihe520/MathModelAgent
MathModelAgent is a Python agent built for math modeling competitions: it analyzes the problem, builds the model, writes and fixes code, and drafts the paper, all in one run. Separate agents handle modeling, coding, and writing, so each stage gets the right model for the job instead of one LLM doing everything.
The output isn't a pile of notes — it's a formatted paper matched to one of 17 contest templates (national and international, Typst-based), backed by a small modeling knowledge base and nine automated checks that catch inconsistent numbers before you submit. Code runs through a local Jupyter interpreter or cloud sandboxes like E2B and daytona, and works with any LLM provider via litellm.
Try the desktop build for a zero-setup start, or run it via Docker, a local Python/Node/Redis install, or as a Claude Code / Codex skill with a single slash command.
https://github.com/jihe520/MathModelAgent
A code review agent that beats Claude Code on precision, at a ninth of the tokens
General-purpose coding agents skimp on large diffs: they skip files, drift on line numbers, and swing wildly with every prompt tweak. Open Code Review fixes that by splitting the job — deterministic engineering decides which files matter and bundles related ones into isolated review units, while an LLM agent does the actual bug hunting with full codebase context.
It comes with a built-in multi-language ruleset for NPE, thread-safety, XSS and SQL injection, and leaves comments precise to the line. It's the same tool that has reviewed code inside Alibaba for two years, across tens of thousands of developers, now open-sourced. Works with any OpenAI- or Anthropic-compatible model endpoint.
Point it at a repo, configure your model of choice, and run
github.com/alibaba/open-code-review
General-purpose coding agents skimp on large diffs: they skip files, drift on line numbers, and swing wildly with every prompt tweak. Open Code Review fixes that by splitting the job — deterministic engineering decides which files matter and bundles related ones into isolated review units, while an LLM agent does the actual bug hunting with full codebase context.
It comes with a built-in multi-language ruleset for NPE, thread-safety, XSS and SQL injection, and leaves comments precise to the line. It's the same tool that has reviewed code inside Alibaba for two years, across tens of thousands of developers, now open-sourced. Works with any OpenAI- or Anthropic-compatible model endpoint.
Point it at a repo, configure your model of choice, and run
ocr on a diff or ocr scan on a whole directory — no fine-tuning, no prompt engineering required.github.com/alibaba/open-code-review