Tailscale had months of reliability problems and eventually traced them into SQLite itself — a bug that had existed for roughly 16 years
https://tailscale.com/blog/sqlite-wal-reset-bug
https://tailscale.com/blog/sqlite-wal-reset-bug
Tailscale
How Tailscale helped find the SQLite WAL-Reset bug
Tailscale and SQLite developers traced maddening corruption incidents to find the WAL-Reset data race, then uncovered a second stale expression index bug.
👍5
A fast, structural YAML diff tool — in a single-dependency binary
https://github.com/szhekpisov/diffyml
https://github.com/szhekpisov/diffyml
GitHub
GitHub - szhekpisov/diffyml: A fast, structural YAML diff tool — in a single-dependency binary
A fast, structural YAML diff tool — in a single-dependency binary - szhekpisov/diffyml
❤5👍3🔥3
Mercado Libre runs its observability platform, O11y events, on ClickHouse Cloud to answer granular, business-level questions like why a specific payment failed.
The team built O11y events to take troubleshooting from days to minutes, with full business-flow visibility and high-cardinality filtering on identifiers like payment and user IDs.
Migrating to ClickHouse Cloud increased query performance by 50x and delivered up to 89% data compression, enabling them to scale from 7 million spans per minute to 400 million and growing in ClickHouse.
https://clickhouse.com/blog/mercado-libre-observability-on-clickhouse-cloud
The team built O11y events to take troubleshooting from days to minutes, with full business-flow visibility and high-cardinality filtering on identifiers like payment and user IDs.
Migrating to ClickHouse Cloud increased query performance by 50x and delivered up to 89% data compression, enabling them to scale from 7 million spans per minute to 400 million and growing in ClickHouse.
https://clickhouse.com/blog/mercado-libre-observability-on-clickhouse-cloud
ClickHouse
How Mercado Libre rebuilt its observability platform on ClickHouse Cloud with 50x faster trace queries | ClickHouse
How Mercado Libre rebuilt its observability platform on ClickHouse Cloud, cutting trace query times from over five minutes to about four seconds (a 50x speedup) with up to 89% compression while ingesting 400 million spans per minute.
❤3👍2🔥1
I always use this service before every deployment. I’ve had 100% uptime ever since: https://deploytarot.com/
Deploy Tarot
The Cards Await — Deploy Tarot
Draw your deployment tarot reading. Pick your role, pick your intent, and let the Major Arcana decide if today is your day.
🤣10
A practical guide to building a local multi-cluster platform engineering lab with vind, Sveltos, and Argo CD. It covers GitOps, label-based deployments, drift correction, and the networking issues you’ll encounter along the way.
https://itnext.io/local-platform-engineering-on-your-laptop-vind-sveltos-and-argocd-2b3e1341ebe7
https://itnext.io/local-platform-engineering-on-your-laptop-vind-sveltos-and-argocd-2b3e1341ebe7
Medium
Local Platform Engineering on Your Laptop | vind, Sveltos and ArgoCD
I’m not new to Kubernetes. Certified my way through most of what the CNCF has to offer. And yet this specific setup, a proper local…
👍3❤1
A Kubernetes operator designed to intelligently manage resource overcommit on pod resource requests.
https://github.com/InditexTech/k8s-overcommit-operator
https://github.com/InditexTech/k8s-overcommit-operator
GitHub
GitHub - InditexTech/k8s-overcommit-operator: A Kubernetes operator designed to intelligently manage resource overcommit on pod…
A Kubernetes operator designed to intelligently manage resource overcommit on pod resource requests. - InditexTech/k8s-overcommit-operator
👍3
Failure is inevitable: Learning from a large outage, and building for reliability in depth at Datadog — Datadog Engineering
https://www.datadoghq.com/blog/engineering/rethinking-reliability/
https://www.datadoghq.com/blog/engineering/rethinking-reliability/
Datadog
Failure is inevitable: Learning from a large outage, and building for reliability in depth at Datadog | Datadog
After a major outage, we re-architected Datadog systems to degrade gracefully under failure. Here’s what we learned—and how we’re building forward.
👍2
Securing every Kubernetes workload at scale — LinkedIn Engineering
https://www.linkedin.com/blog/engineering/infrastructure/securing-every-kubernetes-workload-at-scale
https://www.linkedin.com/blog/engineering/infrastructure/securing-every-kubernetes-workload-at-scale
Linkedin
Securing every Kubernetes workload at scale
👍4
Realtime log viewer with web UI, tail -f for logs with a web interface browser.
https://github.com/logdyhq/logdy-core
https://github.com/logdyhq/logdy-core
GitHub
GitHub - logdyhq/logdy-core: Realtime log viewer with web UI, tail -f for logs with a web interface browser.
Realtime log viewer with web UI, tail -f for logs with a web interface browser. - logdyhq/logdy-core