This media is not supported in your browser
VIEW IN TELEGRAM
π¦ tencent/yolo-master
YOLO-Master: Real-Time Object Detection
Tencent released a new research prototype called YOLO-Master that mixes two powerful ideas to make object detection faster. Instead of using one giant brain for every image, this system picks smaller, specialized experts only when they are needed. This dynamic approach saves computing power while keeping the speed high enough for real-time video. The code is based on official research presented at a major computer vision conference. It shows how splitting tasks between different neural network parts can improve performance without slowing things down. Developers can explore this open-source model to see how modern detection tools are evolving beyond standard single-model designs.
π @hackernewsgithubprojects
YOLO-Master: Real-Time Object Detection
Tencent released a new research prototype called YOLO-Master that mixes two powerful ideas to make object detection faster. Instead of using one giant brain for every image, this system picks smaller, specialized experts only when they are needed. This dynamic approach saves computing power while keeping the speed high enough for real-time video. The code is based on official research presented at a major computer vision conference. It shows how splitting tasks between different neural network parts can improve performance without slowing things down. Developers can explore this open-source model to see how modern detection tools are evolving beyond standard single-model designs.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ jtydhr88/comfyui-ultrashape1
ComfyUI UltraShape 1 Mesh Refinement
Take a rough 3D mesh and instantly sharpen it with ComfyUI UltraShape 1. This plugin adds image-guided refinement to your workflow, turning blurry or blocky geometry into high-quality, detailed models. You feed it a coarse mesh from another tool and a reference photo, and it uses deep learning to fill in the missing details and sharpen the edges. It works directly inside your existing ComfyUI setup, supporting common file formats and offering low-vram modes for smoother performance. This makes upgrading basic 3D assets to professional quality surprisingly simple and accessible for anyone building scenes or characters.
π @hackernewsgithubprojects
ComfyUI UltraShape 1 Mesh Refinement
Take a rough 3D mesh and instantly sharpen it with ComfyUI UltraShape 1. This plugin adds image-guided refinement to your workflow, turning blurry or blocky geometry into high-quality, detailed models. You feed it a coarse mesh from another tool and a reference photo, and it uses deep learning to fill in the missing details and sharpen the edges. It works directly inside your existing ComfyUI setup, supporting common file formats and offering low-vram modes for smoother performance. This makes upgrading basic 3D assets to professional quality surprisingly simple and accessible for anyone building scenes or characters.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ idollab/anytrack
AnyTrack: Tracking Objects with Any Data
Track moving objects in videos using any kind of data input you have. This project unifies visual object tracking by letting models adapt to different modalities like text, audio, or standard images. Instead of being locked into just seeing pixels, AnyTrack accepts various data types to identify and follow targets. It solves the problem of rigid tracking systems by offering flexibility across different sensory inputs. This makes it useful for complex scenes where visual cues alone might fail. The core idea is simple: give it any relevant data, and it finds the object. It is a fascinating step toward more adaptable computer vision tools that work with diverse information sources.
π @hackernewsgithubprojects
AnyTrack: Tracking Objects with Any Data
Track moving objects in videos using any kind of data input you have. This project unifies visual object tracking by letting models adapt to different modalities like text, audio, or standard images. Instead of being locked into just seeing pixels, AnyTrack accepts various data types to identify and follow targets. It solves the problem of rigid tracking systems by offering flexibility across different sensory inputs. This makes it useful for complex scenes where visual cues alone might fail. The core idea is simple: give it any relevant data, and it finds the object. It is a fascinating step toward more adaptable computer vision tools that work with diverse information sources.
π @hackernewsgithubprojects
Media is too big
VIEW IN TELEGRAM
π¦ h-embodvis/simwam
SimWAM: Fast Self-Driving Planning
SimWAM lets you train an autonomous driving planner by borrowing smarts from video generation, then throws the video part away. It co-trains a heavy video model and a lightweight action model, using the videoβs understanding of physics to teach the planner how traffic moves. Once the training is done, you discard the video branch entirely. This leaves a super-fast, self-contained system that predicts driving paths directly from camera images without any heavy lifting. The team even used reinforcement learning to fine-tune the driving style, achieving top-tier results on navigation benchmarks with incredibly low latency. It is a clever shortcut to building efficient self-driving software.
π @hackernewsgithubprojects
SimWAM: Fast Self-Driving Planning
SimWAM lets you train an autonomous driving planner by borrowing smarts from video generation, then throws the video part away. It co-trains a heavy video model and a lightweight action model, using the videoβs understanding of physics to teach the planner how traffic moves. Once the training is done, you discard the video branch entirely. This leaves a super-fast, self-contained system that predicts driving paths directly from camera images without any heavy lifting. The team even used reinforcement learning to fine-tune the driving style, achieving top-tier results on navigation benchmarks with incredibly low latency. It is a clever shortcut to building efficient self-driving software.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ spikelab-jhu/trace-active-reconstruction
Trace Active Reconstruction
trace-active-reconstruction stops robots from wasting time looking at the same spot twice. Instead of looping over already known areas, this project uses a clever planner that actively avoids redundant views. It forces the camera to hunt for fresh, informative geometry until it has fully mapped the environment. The system builds a real-time map and adjusts its path on the fly to maximize what it sees. It is built for simulation but offers a clear window into how smart trajectory planning works. Watch a robot efficiently explore a room in just minutes, proving that smart planning beats random wandering every time.
π @hackernewsgithubprojects
Trace Active Reconstruction
trace-active-reconstruction stops robots from wasting time looking at the same spot twice. Instead of looping over already known areas, this project uses a clever planner that actively avoids redundant views. It forces the camera to hunt for fresh, informative geometry until it has fully mapped the environment. The system builds a real-time map and adjusts its path on the fly to maximize what it sees. It is built for simulation but offers a clear window into how smart trajectory planning works. Watch a robot efficiently explore a room in just minutes, proving that smart planning beats random wandering every time.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ danielmiessler/pai
Pai: Your Personal Growth Engine
Take control of your personal development with Pai, a tool that turns artificial intelligence into a dedicated growth coach. It works by constantly comparing where you are right now against where you actually want to be, then guiding you through a simple step-by-step process to close that gap. Instead of just giving generic advice, it helps you build a personalized system for life and work that adapts as you improve. You can use it to set clear goals, track your progress, and make smarter decisions every single day. Itβs like having a wise mentor who knows your story inside and out, helping you climb higher without getting stuck.
π @hackernewsgithubprojects
Pai: Your Personal Growth Engine
Take control of your personal development with Pai, a tool that turns artificial intelligence into a dedicated growth coach. It works by constantly comparing where you are right now against where you actually want to be, then guiding you through a simple step-by-step process to close that gap. Instead of just giving generic advice, it helps you build a personalized system for life and work that adapts as you improve. You can use it to set clear goals, track your progress, and make smarter decisions every single day. Itβs like having a wise mentor who knows your story inside and out, helping you climb higher without getting stuck.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ maxfreedompollard/compartment
Encrypted Offline AI Memory
Compartment gives your AI agent a permanent, private memory that actually works. Most assistants forget everything the moment a chat ends, but this tool saves every decision, preference, and detail you share, keeping it safe on your own computer. It runs completely offline with zero cloud connection, encrypting every byte including the search indexes so nothing can be stolen or tracked. The surprise? It ships with six thousand pre-loaded facts about computers, codes, and physics, letting agents operate faster without needing the internet at all.
π @hackernewsgithubprojects
Encrypted Offline AI Memory
Compartment gives your AI agent a permanent, private memory that actually works. Most assistants forget everything the moment a chat ends, but this tool saves every decision, preference, and detail you share, keeping it safe on your own computer. It runs completely offline with zero cloud connection, encrypting every byte including the search indexes so nothing can be stolen or tracked. The surprise? It ships with six thousand pre-loaded facts about computers, codes, and physics, letting agents operate faster without needing the internet at all.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ eternityspring/shuohao-skills
Turn Any Novel into a Character Bible
Shuohao Skills is the creative tool that turns a raw novel into a complete character design bible in seconds. You feed it a story and it scans the text to extract every character, merging different names into one profile and backing every detail with exact quotes from the book. The real magic is the auto-generated model sheet for each person, giving you a bust portrait, full body turnaround, and key detail close-ups all in one clean layout. You can even switch the visual style to look like a Studio Ghibli film or stay realistic.
π @hackernewsgithubprojects
Turn Any Novel into a Character Bible
Shuohao Skills is the creative tool that turns a raw novel into a complete character design bible in seconds. You feed it a story and it scans the text to extract every character, merging different names into one profile and backing every detail with exact quotes from the book. The real magic is the auto-generated model sheet for each person, giving you a bust portrait, full body turnaround, and key detail close-ups all in one clean layout. You can even switch the visual style to look like a Studio Ghibli film or stay realistic.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ mco-org/mco
MCO: Compare AI Coding Agents
Run multiple AI coding agents on one task, compare their raw answers side-by-side, and decide which to trust before acting. No guesswork. Just clear, parallel perspectives from tools like Claude, Codex, and Pi. Check it out. Run MCO to orchestrate AI coding agents. Give them the same task, watch them work in parallel, and then compare their actual raw answers side by side. It lets you pick specific tools like Claude or Codex, run them simultaneously, and keep their unedited responses for review. This avoids trusting a single blind spot by letting you see where agents agree or disagree before you commit to code.
π @hackernewsgithubprojects
MCO: Compare AI Coding Agents
Run multiple AI coding agents on one task, compare their raw answers side-by-side, and decide which to trust before acting. No guesswork. Just clear, parallel perspectives from tools like Claude, Codex, and Pi. Check it out. Run MCO to orchestrate AI coding agents. Give them the same task, watch them work in parallel, and then compare their actual raw answers side by side. It lets you pick specific tools like Claude or Codex, run them simultaneously, and keep their unedited responses for review. This avoids trusting a single blind spot by letting you see where agents agree or disagree before you commit to code.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ jonexaiorg/jonex
Jonex: Turn Raw Data into Structured Knowledge
Jonex lets you turn messy documents, videos, and audio into organized, AI-ready knowledge you can actually use. Instead of just searching through files, it builds a smart knowledge graph that understands how different pieces of information connect. You upload your content, and the system parses every type of file to extract facts, then compiles them into a structured ontology. This means when you ask a question, the AI reasons through the connections first, giving you precise answers with clear sources. It is like giving your entire company library a brain, so you get accurate, traceable insights instead of generic summaries.
π @hackernewsgithubprojects
Jonex: Turn Raw Data into Structured Knowledge
Jonex lets you turn messy documents, videos, and audio into organized, AI-ready knowledge you can actually use. Instead of just searching through files, it builds a smart knowledge graph that understands how different pieces of information connect. You upload your content, and the system parses every type of file to extract facts, then compiles them into a structured ontology. This means when you ask a question, the AI reasons through the connections first, giving you precise answers with clear sources. It is like giving your entire company library a brain, so you get accurate, traceable insights instead of generic summaries.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ xiaobiaodu/mobile-gs
Mobile-GS: Real-time 3D on Your Phone
Mobile-GS lets you render detailed three-dimensional scenes directly on your mobile device, turning your phone into a powerful 3D viewer. Usually, this kind of heavy lifting requires expensive desktop computers, but this project shrinks that technology down so it actually works on handhelds. It achieves this by compressing the visual data and optimizing how the graphics card draws it, allowing for smooth, real-time exploration of complex environments without draining your battery or lagging. This is genuinely exciting because it brings professional-grade 3D visualization to devices we already carry every day, opening up new ways to view digital spaces anywhere you go.
π @hackernewsgithubprojects
Mobile-GS: Real-time 3D on Your Phone
Mobile-GS lets you render detailed three-dimensional scenes directly on your mobile device, turning your phone into a powerful 3D viewer. Usually, this kind of heavy lifting requires expensive desktop computers, but this project shrinks that technology down so it actually works on handhelds. It achieves this by compressing the visual data and optimizing how the graphics card draws it, allowing for smooth, real-time exploration of complex environments without draining your battery or lagging. This is genuinely exciting because it brings professional-grade 3D visualization to devices we already carry every day, opening up new ways to view digital spaces anywhere you go.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ nakasyou/local-mcp
local-mcp
local-mcp is a security-first tool that lets artificial intelligence agents edit your computer files without giving them unlimited access. It wraps sensitive operations in a strict sandbox that blocks network access and limits what the agent can do, forcing it to ask for your permission before touching anything important. You keep total control by approving each action through a simple terminal interface, creating a safe bridge between your local projects and automated coding assistants. It turns risky blind automation into a transparent, auditable process you can trust.
π @hackernewsgithubprojects
local-mcp
local-mcp is a security-first tool that lets artificial intelligence agents edit your computer files without giving them unlimited access. It wraps sensitive operations in a strict sandbox that blocks network access and limits what the agent can do, forcing it to ask for your permission before touching anything important. You keep total control by approving each action through a simple terminal interface, creating a safe bridge between your local projects and automated coding assistants. It turns risky blind automation into a transparent, auditable process you can trust.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ oil-oil/codex-deepseek-subagent
Codex DeepSeek Subagent
The Codex DeepSeek Subagent lets you plug DeepSeek directly into OpenAIβs Codex desktop app as a native helper without touching your main model settings. You install a simple skill, hand over your API key once, and the tool verifies the connection by routing a test task and checking the response for a specific confirmation code. It saves your key securely in your operating systemβs credential manager and backs up your existing config so you never lose anything. After that, you just ask Codex to use the DeepSeek helper for text tasks while your main agent handles the heavy lifting.
π @hackernewsgithubprojects
Codex DeepSeek Subagent
The Codex DeepSeek Subagent lets you plug DeepSeek directly into OpenAIβs Codex desktop app as a native helper without touching your main model settings. You install a simple skill, hand over your API key once, and the tool verifies the connection by routing a test task and checking the response for a specific confirmation code. It saves your key securely in your operating systemβs credential manager and backs up your existing config so you never lose anything. After that, you just ask Codex to use the DeepSeek helper for text tasks while your main agent handles the heavy lifting.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ alikon-art/determinflow
DeterminFlow: Deterministic Workflows for Probabilistic AI
DeterminFlow is the workflow runtime that finally brings structure to unpredictable AI models. Most AI tools act like a single overworked assistant who forgets details and costs a fortune as the chat grows longer. DeterminFlow solves this by breaking complex tasks into isolated, versioned nodes where each step only sees the context it needs. This means you get automatic retries, precise cost tracking, and the ability to resume exactly where you left off if something fails. It is already saving developers massive amounts of tokens in real production pipelines. It turns chaotic AI experiments into reliable, shippable services that just work.
π @hackernewsgithubprojects
DeterminFlow: Deterministic Workflows for Probabilistic AI
DeterminFlow is the workflow runtime that finally brings structure to unpredictable AI models. Most AI tools act like a single overworked assistant who forgets details and costs a fortune as the chat grows longer. DeterminFlow solves this by breaking complex tasks into isolated, versioned nodes where each step only sees the context it needs. This means you get automatic retries, precise cost tracking, and the ability to resume exactly where you left off if something fails. It is already saving developers massive amounts of tokens in real production pipelines. It turns chaotic AI experiments into reliable, shippable services that just work.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ codegraphcontext/grapharc
GraphARC: The Safety Gate for AI Agents
GraphARC is the governance layer for AI agents that finally stops wild code execution before it happens. Most tools just watch AI agents make mistakes and hope for the best, but this project installs a strict admission gate between the plan and the run. An artificial intelligence proposes a workflow, and a deterministic checker immediately reviews it against your safety policies and budget. If a step violates your rules, the gate rejects it and refuses to execute, forcing the agent to rewrite its plan. You get a complete audit trail that matches your live dashboard exactly, so you know exactly what cost and what logic was used.
π @hackernewsgithubprojects
GraphARC: The Safety Gate for AI Agents
GraphARC is the governance layer for AI agents that finally stops wild code execution before it happens. Most tools just watch AI agents make mistakes and hope for the best, but this project installs a strict admission gate between the plan and the run. An artificial intelligence proposes a workflow, and a deterministic checker immediately reviews it against your safety policies and budget. If a step violates your rules, the gate rejects it and refuses to execute, forcing the agent to rewrite its plan. You get a complete audit trail that matches your live dashboard exactly, so you know exactly what cost and what logic was used.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ maxteabag/sqlit
Sqlit: The Terminal SQL Client That Feels Like Magic
Sqlit is the terminal app that finally makes database queries fun. It works like a lightweight version of LazyGit but for SQL. You just type sqlit and it pops up a clean, colorful screen right in your terminal. The coolest part is how it spots your running Docker containers automatically. You press Enter, and it connects without you needing to type complex server addresses or ports. It even remembers your passwords safely and lets you use familiar keyboard shortcuts to type code. It is fast, beautiful, and makes looking at data actually enjoyable instead of tedious.
π @hackernewsgithubprojects
Sqlit: The Terminal SQL Client That Feels Like Magic
Sqlit is the terminal app that finally makes database queries fun. It works like a lightweight version of LazyGit but for SQL. You just type sqlit and it pops up a clean, colorful screen right in your terminal. The coolest part is how it spots your running Docker containers automatically. You press Enter, and it connects without you needing to type complex server addresses or ports. It even remembers your passwords safely and lets you use familiar keyboard shortcuts to type code. It is fast, beautiful, and makes looking at data actually enjoyable instead of tedious.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ 0xwilliamortiz/humanizer-cli
Spot AI Text in Your Terminal
Humanizer-cli turns your terminal into a detective for AI writing. It checks your drafts against thirty-three habits that give away language models, like fake enthusiasm and endless em dashes. The tool runs entirely offline with zero dependencies, so your text never leaves your machine. It spots obvious mechanical tells instantly and highlights subtler patterns that need a human reader. You can scan a file, search for specific quirks, or get a ready-to-paste prompt to rewrite the text by hand. It is a clever, lightweight way to ensure your writing sounds like you, not a machine.
π @hackernewsgithubprojects
Spot AI Text in Your Terminal
Humanizer-cli turns your terminal into a detective for AI writing. It checks your drafts against thirty-three habits that give away language models, like fake enthusiasm and endless em dashes. The tool runs entirely offline with zero dependencies, so your text never leaves your machine. It spots obvious mechanical tells instantly and highlights subtler patterns that need a human reader. You can scan a file, search for specific quirks, or get a ready-to-paste prompt to rewrite the text by hand. It is a clever, lightweight way to ensure your writing sounds like you, not a machine.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ ash-2046/suv
SUV: Future Scene Understanding as Video Generation for End-to-End Driving
A single front camera can now predict a car's entire future path by generating four video streams simultaneously. The project called SUV treats driving planning as a video generation task. It looks at what is happening right now and renders a four-second preview of what the road will look like, including lane markings, depth, and moving objects, all in one go. Instead of guessing many possible paths and picking the best one, the system simply watches its own generated video to figure out where to steer. This approach avoids complex selection logic and achieves top scores on autonomous driving benchmarks.
π @hackernewsgithubprojects
SUV: Future Scene Understanding as Video Generation for End-to-End Driving
A single front camera can now predict a car's entire future path by generating four video streams simultaneously. The project called SUV treats driving planning as a video generation task. It looks at what is happening right now and renders a four-second preview of what the road will look like, including lane markings, depth, and moving objects, all in one go. Instead of guessing many possible paths and picking the best one, the system simply watches its own generated video to figure out where to steer. This approach avoids complex selection logic and achieves top scores on autonomous driving benchmarks.
π @hackernewsgithubprojects
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ shrek3onvh5/minimax-h3-nativeaudio-musicvideo-workflow
MiniMax H3 Native-Audio Music Video Workflow
MiniMax H3 Native-Audio Music Video Workflow is the ComfyUI toolkit that locks generated music videos to your original song files instead of letting AI hallucinate new vocals. While many video generators scramble the audio, this project adds a special node that forces the videoβs lip movements to match the exact rhythm and timing of your supplied track, keeping the audio pristine throughout. It also includes a multishot sampler that chains scenes together so the character stays consistent, plus a clever speed boost that predicts video details between steps to render longer clips faster.
π @hackernewsgithubprojects
MiniMax H3 Native-Audio Music Video Workflow
MiniMax H3 Native-Audio Music Video Workflow is the ComfyUI toolkit that locks generated music videos to your original song files instead of letting AI hallucinate new vocals. While many video generators scramble the audio, this project adds a special node that forces the videoβs lip movements to match the exact rhythm and timing of your supplied track, keeping the audio pristine throughout. It also includes a multishot sampler that chains scenes together so the character stays consistent, plus a clever speed boost that predicts video details between steps to render longer clips faster.
π @hackernewsgithubprojects
Media is too big
VIEW IN TELEGRAM
π¦ aut-aisl/queenvis
QueenVIS: Video Segmentation Without Video Training
QueenVIS teaches a computer to track objects across video frames without ever seeing a video. Instead of learning from moving footage, it trains exclusively on static images. It achieves this by adding two simple rules to its object detection queries: predicting where the objectβs center is and matching its visual features to a single snapshot. This trick makes the model so confident about an objectβs look and position that it can guess which object is which in a moving video later on. By removing these rules after training, it relies on a clever memory system to link frames together.
π @hackernewsgithubprojects
QueenVIS: Video Segmentation Without Video Training
QueenVIS teaches a computer to track objects across video frames without ever seeing a video. Instead of learning from moving footage, it trains exclusively on static images. It achieves this by adding two simple rules to its object detection queries: predicting where the objectβs center is and matching its visual features to a single snapshot. This trick makes the model so confident about an objectβs look and position that it can guess which object is which in a moving video later on. By removing these rules after training, it relies on a clever memory system to link frames together.
π @hackernewsgithubprojects