DevOps&SRE Library
19.9K subscribers
431 photos
1 video
2 files
5.52K links
Библиотека статей по теме DevOps и SRE.

Реклама: @ostinostin
Контент: @mxssl

РКН: https://www.gosuslugi.ru/snet/67704b536aa9672b963777b3
Download Telegram
Building a Real k6 Test Suite Against a Live Kubernetes App

In part 1 I covered k6's philosophy and the anatomy of a first test. This post is where things get real — a production-grade test suite running against a live microservices app on a homelab Kubernetes cluster, including what went wrong on the first run and how I debugged it.


https://dev.to/matthew_wimpelberg_79193b/part-2-of-4-building-a-real-k6-test-suite-against-a-live-kubernetes-app-1f81
From Ingress to Gateway API: How We Modernized Networking on Our GKE Cluster

We recently migrated our production GKE cluster from the traditional Ingress controller to the Kubernetes Gateway API — and honestly, we should have done it sooner.


https://the-devops-engineer.medium.com/from-ingress-to-gateway-api-how-we-modernized-networking-on-our-gke-cluster-8409ffb53173
Autoscalable GitLab runners on AWS EC2

This article shares our experience of getting rid of continuously running EC2 instances, setting up scalable GitLab Runners in AWS, and significantly cutting CI infrastructure costs.


https://palark.com/blog/autoscalable-gitlab-runners-on-aws-ec2
IPMan - IPSec Connection Manager for Kubernetes

IPMan is a Kubernetes operator that simplifies the management of IPSec connections, enabling secure communication between your Kubernetes workloads and the outside world.


https://github.com/dialohq/ipman
✨ Как настроить PostgreSQL для продакшен-нагрузок

При переходе к продакшен-нагрузкам важно заранее продумать отказоустойчивость, сценарии переключения между узлами и поведение кластера при плановых работах и сбоях.


29 сентября эксперт Cloud․ru проведет вебинар о том, как обеспечить высокую доступность PostgreSQL в облаке.

В программе:
▶как устроен отказоустойчивый кластер в Evolution Managed PostgreSQL

▶где проходит граница между высокой доступностью и аварийным восстановлением

▶какую роль играет Multi-AZ-архитектура

▶для каких задач используются реплики PostgreSQL

▶как работают ручное и автоматическое переключение

▶что важно учитывать при эксплуатации PostgreSQL в продакшене


Будет интересно всем, кто работает с PostgreSQL и отвечает за доступность баз данных и надежность инфраструктуры.

Зарегистрироваться
Please open Telegram to view this post
VIEW IN TELEGRAM
Klarity

Klarity is an open-source, enterprise-grade Kubernetes observability dashboard built for teams that follow GitOps practices.


https://github.com/selvarajmurugesan90/klarity
kstack

Kstack is a skill pack for Claude Code that helps you perform monitoring, troubleshooting and auditing tasks on your K8s clusters in a smart and efficient way.


https://github.com/kubetail-org/kstack
k8s-overcommit Operator

The k8s-overcommit Operator is a Kubernetes operator designed to intelligently manage resource overcommit on pod resource requests. It automatically adjusts CPU and memory requests based on configurable overcommit classes, enabling better cluster resource utilization while maintaining workload performance.


https://github.com/InditexTech/k8s-overcommit-operator
Heroic saves are near misses

What doesn't usually happen is anyone asking: what if she hadn't been there? Because that heroic save, for all the heartfelt celebration around it, was actually a near miss from a systemic point of view.


https://greatcircle.com/blog/2026/07/21/heroic-saves-are-near-misses
Control and complexity: tension in systems design

In this post I want to discuss how we organize systems by contrasting two families of approaches. The first is about analytical decomposition that aims to maintain control over a system, and the other is based on a perspective of complex systems that resist analysis, which tend to focus on figuring out interactions and mechanisms to foster desirable emergent behaviour.


https://ferd.ca/control-and-complexity-tension-in-systems-design.html
Type Conversion: What Changes (and What Doesn't) When You Start a New SRE Job

Pilots have a specific term for moving from one airplane to another: a "type conversion". Not "learning to fly" again — you already know how to fly. It's the process of taking everything you already know and re-mapping it onto a new machine that does the same job with a different cockpit. Starting a new SRE job is the same thing.


https://billduncan.org/type-conversion
How we tracked down a 16-year-old SQLite bug

At the end of last year, our uptime was pretty shaky. Many of these outages were caused by a single bug, deep in SQLite. It took months of intense forensics to track it down.


https://tailscale.com/blog/sqlite-wal-reset-bug