How Netflix Simplified Batch Compute with Kueue
https://medium.com/netflix-techblog/how-netflix-simplified-batch-compute-with-kueue-87860682629c
As a part of the journey to transition Netflix's compute infrastructure to be more Kubernetes-native, we have leaned into incorporating components from the Kubernetes ecosystem into our container platform Titus. One example of this is our use of Kueue, a cloud-native job queueing system for batch workloads, which has largely replaced the custom queuing and scheduling logic in our homegrown managed batch solution Compute Managed Batch (CMB).
https://medium.com/netflix-techblog/how-netflix-simplified-batch-compute-with-kueue-87860682629c
Server-side apply: what happens when you run kubectl apply
https://learnkube.com/server-side-apply-kubernetes
Server-side apply matters because Kubernetes objects are shared state: it moves field ownership into the API server, so apply-style tools can surface conflicts instead of hiding them as silent overwrites.
https://learnkube.com/server-side-apply-kubernetes
Kubernetes is migrating from SPDY to WebSockets
https://kftray.app/blog/kubernetes-spdy-to-websockets
I maintain an app that builds on top of kubernetes port forwarding, so i track KEP-4006 because the streaming protocol underneath keeps changing. I wrote about this back in april 2024 around Kubernetes 1.30, and six releases later it's changed enough to be worth another look.
https://kftray.app/blog/kubernetes-spdy-to-websockets
Миссия на сегодня:
— обновить платформу, не обновляя весь Kubernetes;
— встроить AppSec, не остановив релизы;
— пережить падение хранилища секретов;
— защитить LLM и понять, куда исчезли все токены;
— усложнить Admission-политики, не положив API-сервер.
Если звучит как особенно бодрое дежурство — приходите 15 октября на DevOps-трек конференции Orion soft «Большая игра». Там эти сценарии разберут по архитектуре, коду и результатам тестов.
📍 Москва
Принять миссию
— обновить платформу, не обновляя весь Kubernetes;
— встроить AppSec, не остановив релизы;
— пережить падение хранилища секретов;
— защитить LLM и понять, куда исчезли все токены;
— усложнить Admission-политики, не положив API-сервер.
Если звучит как особенно бодрое дежурство — приходите 15 октября на DevOps-трек конференции Orion soft «Большая игра». Там эти сценарии разберут по архитектуре, коду и результатам тестов.
📍 Москва
Принять миссию
How I Learned to Stop Worrying and Love the Reconciliation Loop
https://medium.com/@bobbydeveaux/how-i-learned-to-stop-worrying-and-love-the-reconciliation-loop-32928d0d80cf
You've got ten Claude Code agents running. Each is working on a different issue. Agent 3 just created a PR that conflicts with Agent 7's changes. Agent 5 has been stuck on "Analyzing codebase…" for 45 minutes. Welcome to multi-agent chaos. Turns out, scaling AI coding agents from 1 to 10+ is a completely different problem from running a single agent in a terminal.
https://medium.com/@bobbydeveaux/how-i-learned-to-stop-worrying-and-love-the-reconciliation-loop-32928d0d80cf
Securing CI/CD for an open source project: lessons from Cilium
https://cilium.io/blog/2026/05/06/securing-cicd-open-source-lessons-from-cilium
Cilium runs in the kernel-level networking path of millions of Kubernetes pods. If our supply chain were compromised, the blast radius would not be small. Hardening the project against that scenario is something we work on continuously, and we wanted to write down what we actually do, in detail. Most of what follows isn't Cilium-specific: any open source project running CI/CD on GitHub Actions can apply these patterns.
https://cilium.io/blog/2026/05/06/securing-cicd-open-source-lessons-from-cilium
The kubectl 401 That Wasn't a Kubernetes Problem
https://medium.com/@pradhyuman-pandey/the-kubectl-401-that-wasnt-a-kubernetes-problem-4e04438f33ae
A short debugging story about a two-year-old credentials file, an EKS cluster, and the AWS credential chain quietly doing exactly what it was designed to do.
https://medium.com/@pradhyuman-pandey/the-kubectl-401-that-wasnt-a-kubernetes-problem-4e04438f33ae
Building a Kubernetes Raspberry Pi Homelab with K3s
https://medium.com/@chris.allmark/building-a-kubernetes-raspberry-pi-homelab-with-k3s-f934e4b24162
As you'll know if you've ever tried to build a reasonably sized microservices application locally, you can run out of system resources quickly. So, in order to free up some of that valuable memory, I figured I'd offload things to a Kubernetes cluster running on some relatively cheap hardware, and in this post I'll describe the steps I took so that you can do the same.
https://medium.com/@chris.allmark/building-a-kubernetes-raspberry-pi-homelab-with-k3s-f934e4b24162
Architecting Next-Generation Infrastructure: Enterprise Multi-Cluster Management Using Rancher, Virtual Clusters (vCluster), and Kargo GitOps
https://medium.com/@taofeekaoyusuf/architecting-next-generation-infrastructure-enterprise-multi-cluster-management-using-rancher-fd024a400c87
This guide breaks down the ultimate architectural blueprint to solve cluster sprawl and deployment friction for good. By unifying the centralized governance of Rancher, the resource-saving isolation of virtual control planes (vCluster), and the advanced multi-stage pipeline automation of Kargo GitOps, this write-up touches on how to build a highly scalable, zero-touch infrastructure environment.
https://medium.com/@taofeekaoyusuf/architecting-next-generation-infrastructure-enterprise-multi-cluster-management-using-rancher-fd024a400c87
2
Authentication between microservices using Kubernetes identities
https://learnkube.com/microservices-authentication-kubernetes
In this article, we will use Kubernetes Service Accounts and the TokenReview API to authenticate requests between two in-cluster services, then improve the setup with audience-bound projected Service Account tokens.
https://learnkube.com/microservices-authentication-kubernetes
l9gpu: GPU telemetry with workload attribution
https://github.com/last9/gpu-telemetry
DCGM exporter tells you a GPU is hot. It won't tell you whose job is frying it. l9gpu closes the loop. One agent per node emits vendor-neutral OTLP with workload attribution baked in — Kubernetes pod, namespace, deployment; Slurm job, user, partition.
https://github.com/last9/gpu-telemetry
Заглянем внутрь современных инженерных платформ
14 октября Т-Банк приглашает на Platform Engineering Night — вечер для инженеров, которые создают платформы и помогают командам справляться с растущей сложностью разработки.
Что будем делать:
🔴 Обсудим внутренние платформы, которые помогают быстрее разрабатывать, запускать и сопровождать сервисы.
Разберем, как масштабировать платформы, когда растет количество пользователей, сервисов и связей между ними.
🔴 Расскажем про новые Т-ЦОДы — обсудим, как платформы помогают согласовать работу между дата-центрами в разных регионах.
🔴 Покажем, как AI встроен в платформы и весь SDLC — от разработки до эксплуатации.
🔴 Обсудим реальные инженерные кейсы с командами Т-Банка и других технологических компаний.
➡ Когда: 14 октября, начало в 18:00
📍 Где: Онлайн и офлайн в ИТ-хабе Т-Банка в Санкт-Петербурге
Мероприятие бесплатное, торопитесь занять место по ссылке на регистрацию.
14 октября Т-Банк приглашает на Platform Engineering Night — вечер для инженеров, которые создают платформы и помогают командам справляться с растущей сложностью разработки.
Что будем делать:
Разберем, как масштабировать платформы, когда растет количество пользователей, сервисов и связей между ними.
Мероприятие бесплатное, торопитесь занять место по ссылке на регистрацию.
Please open Telegram to view this post
VIEW IN TELEGRAM