DevOps&SRE Library
19.9K subscribers
426 photos
1 video
2 files
5.52K links
Библиотека статей по теме DevOps и SRE.

Реклама: @ostinostin
Контент: @mxssl

РКН: https://www.gosuslugi.ru/snet/67704b536aa9672b963777b3
Download Telegram
kstack

Kstack is a skill pack for Claude Code that helps you perform monitoring, troubleshooting and auditing tasks on your K8s clusters in a smart and efficient way.


https://github.com/kubetail-org/kstack
k8s-overcommit Operator

The k8s-overcommit Operator is a Kubernetes operator designed to intelligently manage resource overcommit on pod resource requests. It automatically adjusts CPU and memory requests based on configurable overcommit classes, enabling better cluster resource utilization while maintaining workload performance.


https://github.com/InditexTech/k8s-overcommit-operator
Heroic saves are near misses

What doesn't usually happen is anyone asking: what if she hadn't been there? Because that heroic save, for all the heartfelt celebration around it, was actually a near miss from a systemic point of view.


https://greatcircle.com/blog/2026/07/21/heroic-saves-are-near-misses
Control and complexity: tension in systems design

In this post I want to discuss how we organize systems by contrasting two families of approaches. The first is about analytical decomposition that aims to maintain control over a system, and the other is based on a perspective of complex systems that resist analysis, which tend to focus on figuring out interactions and mechanisms to foster desirable emergent behaviour.


https://ferd.ca/control-and-complexity-tension-in-systems-design.html
Type Conversion: What Changes (and What Doesn't) When You Start a New SRE Job

Pilots have a specific term for moving from one airplane to another: a "type conversion". Not "learning to fly" again — you already know how to fly. It's the process of taking everything you already know and re-mapping it onto a new machine that does the same job with a different cockpit. Starting a new SRE job is the same thing.


https://billduncan.org/type-conversion
How we tracked down a 16-year-old SQLite bug

At the end of last year, our uptime was pretty shaky. Many of these outages were caused by a single bug, deep in SQLite. It took months of intense forensics to track it down.


https://tailscale.com/blog/sqlite-wal-reset-bug
Optimizing Kubernetes pod deployments for reliability with topology spread constraints

Pod distribution plays a much bigger role in reliability than you might think. By adding a few lines to your manifest, you can ensure your deployments are zone-redundant and evenly scalable. The feature is called topology spread constraints, and in this blog, we'll explain how it works in full detail.


https://www.gremlin.com/blog/optimizing-kubernetes-pod-deployments-for-reliability-with-topology-spread-constraints
rune

Rune is a fast, GPU-accelerated, full-featured IDE and terminal multiplexer, suitable both for automatic and manual programming.


https://github.com/unstablebuild/rune
drop

Linux sandboxing that doesn't get in your way


https://github.com/wrr/drop
gortex

High-performance code-intelligence engine for AI agents and IDE, supports 257 languages, multi repositories, based on graph, with access via CLI, MCP Server, and API. AI coding agents teammate - expose only needed information, cutting token usage up to 50x. 100% local.


https://github.com/zzet/gortex
atlas

Atlas is source control for coding agents. Every agent run produces checkpoints: commits are linked back to the session that made it alongside the prompts, tool calls, and reasoning. You see which agent did exactly what and why.

Run Claude Code, Codex, Atlas's own agent, or anything from the ACP registry side by side against the same codebase, with shared memory so switching agents mid-task doesn't mean starting over.


https://github.com/pacifio/atlas
ax

Declare an agentic task with workspaces and gateway specifications. AX sandboxes it, wires up its workspace, fences its network, and helps running it at scale.

AX is a high-throughput, declarative orchestrator to run billions of autonomous agent workloads in a cluster. It runs on top of Agent Substrate for sandboxed execution and is built to run billions of tasks per cluster. If you have used Kubernetes, ax will feel similar.


https://github.com/google/ax
Библиотека практик: проактивная защита контейнеров

8 октября, 11:00 — стрим «Лаборатории Касперского» для тех, кто пишет и выкатывает сервисы в контейнерах.

Возможности Kaspersky Container Security, обзор контейнерных угроз и экспертная дискуссии о лучших практиках работы с контейнерными инфраструктурами.

Зарегистрироваться
git-bug

Distributed, offline-first bug tracker embedded in git


https://github.com/git-bug/git-bug