DevOps&SRE Library
19.9K subscribers
431 photos
1 video
2 files
5.52K links
Библиотека статей по теме DevOps и SRE.

Реклама: @ostinostin
Контент: @mxssl

РКН: https://www.gosuslugi.ru/snet/67704b536aa9672b963777b3
Download Telegram
Kafka on Kubernetes: Performance Lessons for Any Disk-Heavy Data Service

We recently started migrating Kafka clusters from EC2 to EKS using Strimzi. As soon as we moved the first cluster, we saw persistent disk reads across the brokers and higher latency than we expected on comparable hardware.


https://dev.to/yaakovamar/kafka-on-kubernetes-performance-lessons-for-any-disk-heavy-data-service-3bl5
Your AI just deleted the wrong deployment. Now what?

Picture this. A developer asks an AI assistant to "scale down staging to save costs." The AI, helpful as always, executes: kubectl scale deployment critical-api --replicas=0 -n production. Wrong namespace. Right outcome, wrong cluster. The API is down.


https://medium.com/@mirusser/your-ai-just-deleted-the-wrong-deployment-now-what-d9e3a03bf46c
My Experiments with MCP: Moving Beyond the "Agent Wrapper"

I'm currently working with a client to build out an agent-based automation system designed to reduce the manual labor associated with weekly, monthly, and ad-hoc operational activities.


https://godfreym.medium.com/my-experiments-with-mcp-moving-beyond-the-agent-wrapper-4142bb920f4a
Building a Real k6 Test Suite Against a Live Kubernetes App

In part 1 I covered k6's philosophy and the anatomy of a first test. This post is where things get real — a production-grade test suite running against a live microservices app on a homelab Kubernetes cluster, including what went wrong on the first run and how I debugged it.


https://dev.to/matthew_wimpelberg_79193b/part-2-of-4-building-a-real-k6-test-suite-against-a-live-kubernetes-app-1f81
From Ingress to Gateway API: How We Modernized Networking on Our GKE Cluster

We recently migrated our production GKE cluster from the traditional Ingress controller to the Kubernetes Gateway API — and honestly, we should have done it sooner.


https://the-devops-engineer.medium.com/from-ingress-to-gateway-api-how-we-modernized-networking-on-our-gke-cluster-8409ffb53173
Autoscalable GitLab runners on AWS EC2

This article shares our experience of getting rid of continuously running EC2 instances, setting up scalable GitLab Runners in AWS, and significantly cutting CI infrastructure costs.


https://palark.com/blog/autoscalable-gitlab-runners-on-aws-ec2
IPMan - IPSec Connection Manager for Kubernetes

IPMan is a Kubernetes operator that simplifies the management of IPSec connections, enabling secure communication between your Kubernetes workloads and the outside world.


https://github.com/dialohq/ipman
✨ Как настроить PostgreSQL для продакшен-нагрузок

При переходе к продакшен-нагрузкам важно заранее продумать отказоустойчивость, сценарии переключения между узлами и поведение кластера при плановых работах и сбоях.


29 сентября эксперт Cloud․ru проведет вебинар о том, как обеспечить высокую доступность PostgreSQL в облаке.

В программе:
▶как устроен отказоустойчивый кластер в Evolution Managed PostgreSQL

▶где проходит граница между высокой доступностью и аварийным восстановлением

▶какую роль играет Multi-AZ-архитектура

▶для каких задач используются реплики PostgreSQL

▶как работают ручное и автоматическое переключение

▶что важно учитывать при эксплуатации PostgreSQL в продакшене


Будет интересно всем, кто работает с PostgreSQL и отвечает за доступность баз данных и надежность инфраструктуры.

Зарегистрироваться
Please open Telegram to view this post
VIEW IN TELEGRAM
Klarity

Klarity is an open-source, enterprise-grade Kubernetes observability dashboard built for teams that follow GitOps practices.


https://github.com/selvarajmurugesan90/klarity
kstack

Kstack is a skill pack for Claude Code that helps you perform monitoring, troubleshooting and auditing tasks on your K8s clusters in a smart and efficient way.


https://github.com/kubetail-org/kstack
k8s-overcommit Operator

The k8s-overcommit Operator is a Kubernetes operator designed to intelligently manage resource overcommit on pod resource requests. It automatically adjusts CPU and memory requests based on configurable overcommit classes, enabling better cluster resource utilization while maintaining workload performance.


https://github.com/InditexTech/k8s-overcommit-operator
Heroic saves are near misses

What doesn't usually happen is anyone asking: what if she hadn't been there? Because that heroic save, for all the heartfelt celebration around it, was actually a near miss from a systemic point of view.


https://greatcircle.com/blog/2026/07/21/heroic-saves-are-near-misses