The kubectl 401 That Wasn't a Kubernetes Problem
https://medium.com/@pradhyuman-pandey/the-kubectl-401-that-wasnt-a-kubernetes-problem-4e04438f33ae
A short debugging story about a two-year-old credentials file, an EKS cluster, and the AWS credential chain quietly doing exactly what it was designed to do.
https://medium.com/@pradhyuman-pandey/the-kubectl-401-that-wasnt-a-kubernetes-problem-4e04438f33ae
Building a Kubernetes Raspberry Pi Homelab with K3s
https://medium.com/@chris.allmark/building-a-kubernetes-raspberry-pi-homelab-with-k3s-f934e4b24162
As you'll know if you've ever tried to build a reasonably sized microservices application locally, you can run out of system resources quickly. So, in order to free up some of that valuable memory, I figured I'd offload things to a Kubernetes cluster running on some relatively cheap hardware, and in this post I'll describe the steps I took so that you can do the same.
https://medium.com/@chris.allmark/building-a-kubernetes-raspberry-pi-homelab-with-k3s-f934e4b24162
Architecting Next-Generation Infrastructure: Enterprise Multi-Cluster Management Using Rancher, Virtual Clusters (vCluster), and Kargo GitOps
https://medium.com/@taofeekaoyusuf/architecting-next-generation-infrastructure-enterprise-multi-cluster-management-using-rancher-fd024a400c87
This guide breaks down the ultimate architectural blueprint to solve cluster sprawl and deployment friction for good. By unifying the centralized governance of Rancher, the resource-saving isolation of virtual control planes (vCluster), and the advanced multi-stage pipeline automation of Kargo GitOps, this write-up touches on how to build a highly scalable, zero-touch infrastructure environment.
https://medium.com/@taofeekaoyusuf/architecting-next-generation-infrastructure-enterprise-multi-cluster-management-using-rancher-fd024a400c87
2
Authentication between microservices using Kubernetes identities
https://learnkube.com/microservices-authentication-kubernetes
In this article, we will use Kubernetes Service Accounts and the TokenReview API to authenticate requests between two in-cluster services, then improve the setup with audience-bound projected Service Account tokens.
https://learnkube.com/microservices-authentication-kubernetes
l9gpu: GPU telemetry with workload attribution
https://github.com/last9/gpu-telemetry
DCGM exporter tells you a GPU is hot. It won't tell you whose job is frying it. l9gpu closes the loop. One agent per node emits vendor-neutral OTLP with workload attribution baked in — Kubernetes pod, namespace, deployment; Slurm job, user, partition.
https://github.com/last9/gpu-telemetry
Заглянем внутрь современных инженерных платформ
14 октября Т-Банк приглашает на Platform Engineering Night — вечер для инженеров, которые создают платформы и помогают командам справляться с растущей сложностью разработки.
Что будем делать:
🔴 Обсудим внутренние платформы, которые помогают быстрее разрабатывать, запускать и сопровождать сервисы.
Разберем, как масштабировать платформы, когда растет количество пользователей, сервисов и связей между ними.
🔴 Расскажем про новые Т-ЦОДы — обсудим, как платформы помогают согласовать работу между дата-центрами в разных регионах.
🔴 Покажем, как AI встроен в платформы и весь SDLC — от разработки до эксплуатации.
🔴 Обсудим реальные инженерные кейсы с командами Т-Банка и других технологических компаний.
➡ Когда: 14 октября, начало в 18:00
📍 Где: Онлайн и офлайн в ИТ-хабе Т-Банка в Санкт-Петербурге
Мероприятие бесплатное, торопитесь занять место по ссылке на регистрацию.
14 октября Т-Банк приглашает на Platform Engineering Night — вечер для инженеров, которые создают платформы и помогают командам справляться с растущей сложностью разработки.
Что будем делать:
Разберем, как масштабировать платформы, когда растет количество пользователей, сервисов и связей между ними.
Мероприятие бесплатное, торопитесь занять место по ссылке на регистрацию.
Please open Telegram to view this post
VIEW IN TELEGRAM
ArgoCD Mobile
https://github.com/argoproj-labs/mobile-for-argocd
A native iOS and Android app for monitoring and managing Argo CD deployments from your phone.
https://github.com/argoproj-labs/mobile-for-argocd
2
Этот пост видят только приглашенные на конференцию
21 октября для вас и ваших коллег Т-Банк проводит онлайн-встречу «SRE-техтолк». Приглашают SRE, DevOps-инженеров и разработчиков.
Практикующие эксперты приготовили классную программу с реальными кейсами и без абстракций. Вас ждет:
Разбор задач Т-Банка и эффективных практик.
Подходы к надежности, о которых редко рассказывают публично.
Адаптация инструментов под разные задачи.
Приходите послушать трех специалистов с разносторонним опытом, обсудить ваши задачи и погрузиться в бигтех.
Встреча пройдет онлайн. Участие бесплатное, а регистрация — по ссылке!
21 октября для вас и ваших коллег Т-Банк проводит онлайн-встречу «SRE-техтолк». Приглашают SRE, DevOps-инженеров и разработчиков.
Практикующие эксперты приготовили классную программу с реальными кейсами и без абстракций. Вас ждет:
Разбор задач Т-Банка и эффективных практик.
Подходы к надежности, о которых редко рассказывают публично.
Адаптация инструментов под разные задачи.
Приходите послушать трех специалистов с разносторонним опытом, обсудить ваши задачи и погрузиться в бигтех.
Встреча пройдет онлайн. Участие бесплатное, а регистрация — по ссылке!
Ballast: Kubernetes right-sizing operator
https://github.com/Tight-Line/ballast
Ballast is a Kubernetes operator that automatically right-sizes workload resource requests and limits based on real operational history. It is a more active alternative to Fairwinds Goldilocks: rather than suggesting changes, it applies them — at admission time and on running pods via in-place resize (Kubernetes 1.35+).
https://github.com/Tight-Line/ballast
Cordium - Kubernetes sandboxes with secretless access
https://github.com/octelium/cordium
Cordium is a free and open source, self-hosted, identity-based sandbox platform built on Kubernetes and Octelium. Cordium is a general-purpose platform that provides isolated, reproducible isolated sandboxes for developers, AI agents, and automated workloads.
https://github.com/octelium/cordium
Benchmarking LLM Inference with Production Agent Traces
https://medium.com/inference-perf/benchmarking-llm-inference-with-production-agent-traces-f47f7f994aff
Most load-testing tools, however, were designed for a simpler world. They assume requests are independent, stateless, and interchangeable. But an AI agent doesn't work that way. This gap was the motivation behind OpenTelemetry (OTel) Trace Replay, a new capability in Inference Perf, the Kubernetes SIG tool for GenAI inference benchmarking.
https://medium.com/inference-perf/benchmarking-llm-inference-with-production-agent-traces-f47f7f994aff