đź’» The byte is not enough
291 subscribers
1 photo
10 files
64 links
Personal hand-picked collection of articles and tutorials on the matter of software engineering and computer science. Occasionally on science, history, or linguistics.

@virtyaluk for any inquiries.

https://modern-dev.com/
https://github.com/virtyaluk
Download Telegram
zab.totally-ordered-broadcast-protocol.2008.pdf
264.7 KB
🧑‍🔬 Architecture of ZAB – ZooKeeper Atomic Broadcast protocol

The ZAB protocol ensures that the Zookeeper replication is done in order and is also responsible for the election of leader nodes and the restoration of any failed nodes. In a Zookeeper ecosystem, the leader node is the heart of everything; every cluster has one leader node and the rest of the nodes are followers. All incoming client requests and state changes are received at first by the leader with responsibility to replicate it across all its followers (and itself). All incoming read requests are also load balanced by the leader within itself and its followers.

Original ZAB paper:
https://marcoserafini.github.io/papers/zab.pdf

Implementation details of ZAB:
http://www.tcs.hut.fi/Studies/T-79.5001/reports/2012-deSouzaMedeiros.pdf

#systemdesign #zab #consensus #total #order #broadcast #design #quorum #performance #faulttolerance #yahoo #paper #research #scalability

đź’» The byte is not enough
Building Facebook’s service encryption infrastructure

What does it take to incorporate security protocols into a system of millions of services running on thousands of machines across the world? You could use some well-known technologies as Kerberos, but will soon realize it does not work on a big scale. You would then probably stick to the idea of building a custom in-house solution, which is exactly what Facebook did to suit their needs in security and operability without compromising performance.

https://engineering.fb.com/2019/05/29/security/service-encryption/

#systemdesign #security #facebook #kerberos #tls #encryption #traffic #networking #dns #performance #faulttolerance #scalability #infrastructure

đź’» The byte is not enough
🧑‍🔬 CORFU: A Distributed Shared Log

Despite almost forty years of research into replicated storage schemes, the only approach so far to scale up capacity and throughput has been to shard data and trade consistency for performance. The CORFU system breaks this seeming tradeoff by organizing a cluster of drives as a single, shared log. CORFU offers a single-copy semantics at cluster-scale speeds, providing a scalable source of atomicity and durability for distributed systems.

https://www.youtube.com/watch?v=GmVQVT9aZfU

Original paper:
https://www.cs.utexas.edu/~lorenzo/corsi/cs380d/papers/a10-balakrishnan.pdf

#systemdesign #corfu #microsoft #design #reliability #algorithms #distributed #log #replication #consensus #performance #scalability #faulttolerance #paper #research #infrastructure

đź’» The byte is not enough
🧑‍🔬 Delos: Simple, flexible storage for the Facebook control plane

Facebook Delos — is a fundamentally new architecture for building replicated storage systems. Its modular, layered design provides flexibility and simplicity without sacrificing performance or reliability. Delos enables fast time-to-deployment for new storage systems — Facebook deployed an initial version in production within eight months — as well as safe, rapid evolution. Facebook swapped in a new ordering mechanism to obtain 10x lower latency without any service downtime.

https://www.youtube.com/watch?v=wd-GC_XhA2g&t=314s

Blog post presentation:
https://engineering.fb.com/2019/06/06/data-center-engineering/delos/

Original paper:
https://www.usenix.org/system/files/osdi20-balakrishnan.pdf

#systemdesign #delos #facebook #distributed #shared #log #algorithms #design #consensus #performance #scalability #faulttolerance #paper #research #infrastructure #storage #controlplane #corfu #paxos #zookeeper #zab

đź’» The byte is not enough
Fighting spam with Guardian, a real-time analytics and rules engine

Another great story of a practical approach to fight spam at big scale that is being applied at Pinterest. Though this isn't a story of inventing some tremendous and big all-in-one solution, it still provides a nice perspective on how to grow from a small and clumsy solution to a big and robust killer feature.

https://medium.com/pinterest-engineering/fighting-spam-with-guardian-a-real-time-analytics-and-rules-engine-938e7e61fa27

#systemdesign #infrastructure #pinterest #design #scalability #performance #bigdata #hive #kafka #presto #guardian #spam #realtime #analytics #safety #security

đź’» The byte is not enough
(Strong) Eventual Consistency and Conflict-free Replicated Data Types (CRDT)

Conflict-free Replicated Data Types (CRDTs) are an increasingly popular family of algorithms for optimistic replication. They allow data to be concurrently updated on several replicas, even while those replicas are offline, and provide a robust way of merging those updates back into a consistent state. CRDTs are used in geo-replicated databases, multi-user collaboration software, distributed processing frameworks, and various other systems.

However, while the basic principles of CRDTs are now quite well known, many challenging problems are lurking below the surface. It turns out that CRDTs are easy to implement badly. Many published algorithms have anomalies that cause them to behave strangely in some situations. Simple implementations often have terrible performance, and making the performance good is challenging.

Martin Kleppmann's talk on CRDT at Hydra Conf:
https://www.youtube.com/watch?v=PMVBuMK_pJY

Sean Cribbs discusses Convergent Replicated Data Types, data structures that tolerate eventual consistency:
https://www.infoq.com/presentations/CRDT/

Marc Shapiro presentation on SEC and CRTD at Microsoft Research:
https://www.youtube.com/watch?v=oyUHd894w18

#systemdesign #algorithms #sec #strong #consistency #crdt #replication #consensus #conflicts #resolution #riak #math #distributed

đź’» The byte is not enough
đź•° Blast from the past: Please stop calling databases CP or AP

A post from 2015 where Martin Kleppmann argues about usage of CAP-theorem in context of modern distributed systems. Martin does a great job demonstrating that people tend to attribute terms like consistency and availability with the meaning that differ from the one given in the original Brewer's CAP Theorem paper, and as thus demonstrates that almost no modern distributed system could be characterized with "A" or "C" property.

https://martin.kleppmann.com/2015/05/11/please-stop-calling-databases-cp-or-ap.html

#systemdesign #theory #research #cap #brewer #kleppmann #consistency #availability #partitiontolerance #zab #consensus #linearizability #zookeeper #riak #dynamodb #cassandra #voldemort #cp #aperture

đź’» The byte is not enough
6 Things You Need to Know About Kafka Before Using it in a System Design Interview

In a typical system design round, there's a chance you may want to leverage Kafka capabilities in your designated system. But what if an interviewer start asking you questions about Kafka's inner mechanics? Will you be able to answer those questions? Here is a list of 6 most used/interesting aspects of Kafka, knowing which will greatly improve your overall system design skills.

https://levelup.gitconnected.com/6-things-you-need-to-know-about-kafka-before-using-it-in-a-system-design-interview-1fc31451732c

#systemdesign #kafka #distributed #queue #log #infrastructure #design #shared #commit #consumer #producer #faulttolerance #reliability #scalability #performance #broker #zookeeper

đź’» The byte is not enough
How Google Spanner Assigns Commit Timestamps — The Secret Sauce of Its Strong Consistency

Spanner is Google’s global scale and synchronously replicated relational database. Its research paper was first published in 2012. Spanner received much acclaim because it was the first system to distribute data at global scale and support strongly consistent distributed transactions. That seems to break the CAP theorem. But in fact, Spanner chooses consistency over availability in face of network partition. It’s just that Google’s data center infrastructure is so reliable that external users typically don’t worry about its outages. Google now offers Spanner as a service through its cloud platform.

https://levelup.gitconnected.com/how-google-spanner-assigns-commit-timestamps-the-secret-sauce-of-its-strong-consistency-8bc143614f26

#systemdesign #google #spanner #transactions #commit #distributed #database #infrastructure #scalability #consistency #quorum #performance #replication #truetime #api #clocks #synchronization #atomicity

đź’» The byte is not enough
The System Design Ideas Behind Advanced Search Functions

Search, in its most basic form, can be implemented as a plain inverted index. We build the index as we add/delete/update documents, and use it during search to go from a search term to the documents that contain it. That’s all good. But if you’re more ambitious than building a simplistic lookup by keyword system, read on. This blog post is a gentle introduction to the core ideas behind some of the advanced search functions. By the way, you may recognize that a lot of the ideas here are based on the lucene index.

https://levelup.gitconnected.com/the-system-design-ideas-behind-advanced-search-functions-fa9c7d9010a3

#algorithms #search #design #inverted #index #lucene #scoring #ranking #faceting #mutation

đź’» The byte is not enough
A Tricky System Design Interview Question: Explain Server Monitoring

A proactive view on how to approach server monitoring during the system design interview rounds. The article is full of unnecessary details, but it still provides a good point on topics like metrics aggregation, time-series databases and pull/push model trade-offs.

https://betterprogramming.pub/a-tricky-system-design-interview-question-explain-server-monitoring-c5be0ce54a30

#systemdesign #infrastructure #sdi #monitoring #metrics #timeseries #aggregation #prometheus #gorilla #scalability #faulttolerance #alerting

đź’» The byte is not enough
Building an Append-only Log From Scratch

Append-only log, sometimes called write ahead log, is a fundamental building block in all modem databases. It’s used to persist mutation commands for recovery purposes. A command to change the database’s state will first be recorded in such a log before applying to the database. Should the database crash and lose updates, its state can be restored by applying the commands since the most recent snapshot backup. The log’s append-only nature is partly due to functional design — because we only ever need to add new commands to the end — and partly due to performance need — sequential file access is orders of magnitude faster than random file access.

https://eileen-code4fun.medium.com/building-an-append-only-log-from-scratch-e8712b49c924

#design #wal #append #log #distributed #shared #commit #log #database #level #leveldb #compaction #segments #lsmtree

đź’» The byte is not enough
Log Structured Merge Tree (LSM-tree) Implementations (a Demo and LevelDB)

Log structured merge tree, or LSM-tree, is a famous data structure that has been widely adopted by many modern “big data” products, such as BigTable, HBase, LevelDB, etc. Its core idea is very simple and perhaps somewhat counterintuitive if you’re used to the traditional database architecture.

https://eileen-code4fun.medium.com/log-structured-merge-tree-lsm-tree-implementations-a-demo-and-leveldb-d5e028257330

#design #systemdesign #database #lsmtree #leveldb #sstable #index #wal #faulttolerance #performance #scalability

đź’» The byte is not enough
System Design Interview: Distributed Top K Frequent Elements in Stream

The top k frequent elements question is a classic coding interview question. Viewing through the lens of system design interview however, it suddenly has a special appeal. But before we start our usual system design interview discussion, let’s quickly talk about how it’s solved in the coding interview context, which will facilitate the design discussion that follows.

https://levelup.gitconnected.com/system-design-interview-distributed-top-k-frequent-elements-in-stream-2e92d63d777e

#systemdesign #algorithms #design #topK #frequency #distributed #mapreduce #streaming #kafka #scalability #performance #countminsketch

đź’» The byte is not enough