consistent_hashing_and_random_trees_distributed_caching_protocols.pdf
179.9 KB
🧑‍🔬 Consistent Hashing and Random Trees: Distributed Caching Protocols for Relieving Hot Spots on the World Wide Web
Another fundamental work on distributed hashing protocols that influenced distributed systems' evolution. Nowadays, consistent hashing is being used in many software distributions like Amazon Dynamo, Apache Cassandra, Riak, Voldemort to name a few. Few big online platforms are known to implement consistent hashing algorithms to scale for performance, availability, and reliability.
More concise take on the matter:
https://www.toptal.com/big-data/consistent-hashing
#systemsdesign #distributed #design #reliability #performance #availability #scalability #research #paper #consistent #hashing
đź’» The byte is not enough
Another fundamental work on distributed hashing protocols that influenced distributed systems' evolution. Nowadays, consistent hashing is being used in many software distributions like Amazon Dynamo, Apache Cassandra, Riak, Voldemort to name a few. Few big online platforms are known to implement consistent hashing algorithms to scale for performance, availability, and reliability.
More concise take on the matter:
https://www.toptal.com/big-data/consistent-hashing
#systemsdesign #distributed #design #reliability #performance #availability #scalability #research #paper #consistent #hashing
đź’» The byte is not enough
Twine_A_Unified_Cluster_Management_System_for_Shared_Infrastructure.pdf
1.1 MB
🧑‍🔬 Twine: A Unified Cluster Management System for Shared Infrastructure
If you as me were impressed by the impressive work Google done in their Borg system, and it's successor Kubernetes, then you will definitely enjoy learning how Facebook makes use of their infrastructure in an astonishing paper on Facebook Twine. This tremendous work benefits from experience gained through developing and managing other popular systems like Borg, Kubernetes or Mesos, and aims to scale to more than a million of machines.
https://engineering.fb.com/2019/06/06/data-center-engineering/twine/
#systemdesign #facebook #twine #distributed #orchestrator #design #reliability #performance #scalability #faulttolerance #research #paper #borg #kubernetes #mesos
đź’» The byte is not enough
If you as me were impressed by the impressive work Google done in their Borg system, and it's successor Kubernetes, then you will definitely enjoy learning how Facebook makes use of their infrastructure in an astonishing paper on Facebook Twine. This tremendous work benefits from experience gained through developing and managing other popular systems like Borg, Kubernetes or Mesos, and aims to scale to more than a million of machines.
https://engineering.fb.com/2019/06/06/data-center-engineering/twine/
#systemdesign #facebook #twine #distributed #orchestrator #design #reliability #performance #scalability #faulttolerance #research #paper #borg #kubernetes #mesos
đź’» The byte is not enough
nsdi13-final170_update.pdf
370.1 KB
🧑‍🔬 Scaling Memcache at Facebook
While reading Facebook's Twine cluster management system white paper, I noticed an interesting thing in section 5 where paper authors claim to have a highly optimized memcached deployment that can handle "930k lookups per second on an 18-core/36-hyperthread machine". Just think for a second, 930 000 lookup requests per second on a single machine. How the heck they could achieve this kind of performance. To answer this and many other questions, I went on reading another Facebook white paper on optimizing Memcached for meeting the world's largest social network needs.
A short version listing all the big things Facebook incorporated into their version of Memcached can be found here:
https://medium.com/@shagun/scaling-memcache-at-facebook-1ba77d71c082
#systemdesign #facebook #memcached #distributed #caching #design #reliability #performance #scalability #faulttolerance #research #paper #redis #dht
đź’» The byte is not enough
While reading Facebook's Twine cluster management system white paper, I noticed an interesting thing in section 5 where paper authors claim to have a highly optimized memcached deployment that can handle "930k lookups per second on an 18-core/36-hyperthread machine". Just think for a second, 930 000 lookup requests per second on a single machine. How the heck they could achieve this kind of performance. To answer this and many other questions, I went on reading another Facebook white paper on optimizing Memcached for meeting the world's largest social network needs.
A short version listing all the big things Facebook incorporated into their version of Memcached can be found here:
https://medium.com/@shagun/scaling-memcache-at-facebook-1ba77d71c082
#systemdesign #facebook #memcached #distributed #caching #design #reliability #performance #scalability #faulttolerance #research #paper #redis #dht
đź’» The byte is not enough
atc13-bronson.pdf
729.7 KB
🧑‍🔬 TAO: Facebook’s Distributed Data Store for the Social Graph
Have you ever wondered how Facebook manages its social graph, containing petabytes of user data? What techniques do they apply to serve billions of reads and millions of writes each second? All this and much more in another great white paper on Facebook's TAO - a geographically distributed data store that provides efficient and timely access to the social graph for Facebook’s demanding workload using a fixed set of queries.
https://engineering.fb.com/2013/06/25/core-data/tao-the-power-of-the-graph/
USENIX ATC '13 - TAO: Facebook’s Distributed Data Store for the Social Graph:
https://www.youtube.com/watch?v=sNIvHttFjdI
#systemdesign #facebook #memcached #tao #socialgraph #api #caching #design #reliability #performance #scalability #faulttolerance #consistency #research #paper #mysql #infrastructure
đź’» The byte is not enough
Have you ever wondered how Facebook manages its social graph, containing petabytes of user data? What techniques do they apply to serve billions of reads and millions of writes each second? All this and much more in another great white paper on Facebook's TAO - a geographically distributed data store that provides efficient and timely access to the social graph for Facebook’s demanding workload using a fixed set of queries.
https://engineering.fb.com/2013/06/25/core-data/tao-the-power-of-the-graph/
USENIX ATC '13 - TAO: Facebook’s Distributed Data Store for the Social Graph:
https://www.youtube.com/watch?v=sNIvHttFjdI
#systemdesign #facebook #memcached #tao #socialgraph #api #caching #design #reliability #performance #scalability #faulttolerance #consistency #research #paper #mysql #infrastructure
đź’» The byte is not enough
paxos-simple.pdf
92.8 KB
🧑‍🔬 Paxos Made Simple
In the year 1989 Leslie Lamport, a known computer scientist in the field of distributed systems, published his tremendous work on Paxos — a family of protocols for solving consensus in a network of unreliable or fallible processors. Though from the very beginning, the proposed algorithm was diminished by computer science society due to its complexity, it started gaining significant recognition after almost 10 years since first published having a second coming in 1998. In late 2001, Lamport published a simplified version of the original paper, discarding unnecessary information and providing the description on the backbone of Paxos protocol.
Paxos Simplified by Chris Colohan:
https://www.youtube.com/watch?v=SRsK-ZXTeZ0
#systemdesign #paxos #consensus #lamport #design #reliability #performance #faulttolerance #scalability #consistency #quorum #research #paper #infrastructure #distributed
đź’» The byte is not enough
In the year 1989 Leslie Lamport, a known computer scientist in the field of distributed systems, published his tremendous work on Paxos — a family of protocols for solving consensus in a network of unreliable or fallible processors. Though from the very beginning, the proposed algorithm was diminished by computer science society due to its complexity, it started gaining significant recognition after almost 10 years since first published having a second coming in 1998. In late 2001, Lamport published a simplified version of the original paper, discarding unnecessary information and providing the description on the backbone of Paxos protocol.
Paxos Simplified by Chris Colohan:
https://www.youtube.com/watch?v=SRsK-ZXTeZ0
#systemdesign #paxos #consensus #lamport #design #reliability #performance #faulttolerance #scalability #consistency #quorum #research #paper #infrastructure #distributed
đź’» The byte is not enough
16cb30b4b92fd4989b8619a61752a2387c6dd474.pdf
186.2 KB
🧑‍🔬 MapReduce: Simplified Data Processing on Large Clusters
Another seminal work from the past that established distributed systems' evolution for decades ahead. The MapReduce model is probably the most well know programming model designed for processing and generating big data sets with a parallel, distributed algorithm on a cluster.
#systemdesign #mapreduce #design #performance #faulttolerance #scalability #research #paper #infrastructure #distributed #computation #cluster #gfs
đź’» The byte is not enough
Another seminal work from the past that established distributed systems' evolution for decades ahead. The MapReduce model is probably the most well know programming model designed for processing and generating big data sets with a parallel, distributed algorithm on a cluster.
#systemdesign #mapreduce #design #performance #faulttolerance #scalability #research #paper #infrastructure #distributed #computation #cluster #gfs
đź’» The byte is not enough
Turbine_Facebook’s_Service_Management_Platform_for_Stream_Processing.pdf
1.9 MB
🧑‍🔬 Turbine: Facebook’s Service Management Platform
for Stream Processing
A scalable service management platform for Facebook’s stream processing service. Turbine is designed to bridge the gap between the capabilities of existing general-purpose cluster management frameworks like Tupperware and Facebook’s stream processing requirements. In production for several years now, Turbine has enabled a boom in stream processing at Facebook.
https://engineering.fb.com/2020/04/21/data-infrastructure/turbine/
#systemdesign #turbine #facebook #streaming #processing #design #performance #acid #reliability #faulttolerance #scalability #research #paper #infrastructure #distributed #cluster #management
đź’» The byte is not enough
for Stream Processing
A scalable service management platform for Facebook’s stream processing service. Turbine is designed to bridge the gap between the capabilities of existing general-purpose cluster management frameworks like Tupperware and Facebook’s stream processing requirements. In production for several years now, Turbine has enabled a boom in stream processing at Facebook.
https://engineering.fb.com/2020/04/21/data-infrastructure/turbine/
#systemdesign #turbine #facebook #streaming #processing #design #performance #acid #reliability #faulttolerance #scalability #research #paper #infrastructure #distributed #cluster #management
đź’» The byte is not enough
👍1
LogDevice: a distributed data store for logs
A log is the simplest way to record an ordered sequence of immutable records and store them reliably. Build a data intensive distributed service and chances are you will need a log or two somewhere. At Facebook, we build a lot of big distributed services that store and process data. Want to connect two stages of a data processing pipeline without having to worry about flow control or data loss? Have one stage write into a log and the other read from it. Maintaining an index on a large distributed database? Have the indexing service read the update log to apply all the changes in the right order. Got a sequence of work items to be executed in a specific order a week later? Write them into a log, have the consumer lag a week. Dream of distributed transactions? A log with enough capacity to order all your writes makes them possible. Durability concerns? Use a write-ahead log.
https://engineering.fb.com/2017/08/31/core-data/logdevice-a-distributed-data-store-for-logs/
#systemdesign #logdevice #log #wal #logsdb #rocksdb #consensus #paxos #quorum #lsmtree #facebook #storage #design #performance #scalability #research #faulttolerance #infrastructure
đź’» The byte is not enough
A log is the simplest way to record an ordered sequence of immutable records and store them reliably. Build a data intensive distributed service and chances are you will need a log or two somewhere. At Facebook, we build a lot of big distributed services that store and process data. Want to connect two stages of a data processing pipeline without having to worry about flow control or data loss? Have one stage write into a log and the other read from it. Maintaining an index on a large distributed database? Have the indexing service read the update log to apply all the changes in the right order. Got a sequence of work items to be executed in a specific order a week later? Write them into a log, have the consumer lag a week. Dream of distributed transactions? A log with enough capacity to order all your writes makes them possible. Durability concerns? Use a write-ahead log.
https://engineering.fb.com/2017/08/31/core-data/logdevice-a-distributed-data-store-for-logs/
#systemdesign #logdevice #log #wal #logsdb #rocksdb #consensus #paxos #quorum #lsmtree #facebook #storage #design #performance #scalability #research #faulttolerance #infrastructure
đź’» The byte is not enough
Engineering at Meta
LogDevice: a distributed data store for logs
Visit the post for more.
Database Storage Engines: B-Tree vs LSM-Tree
Have you ever concern yourself with the question of how modern database systems' internal storage works? Well, there are two popular ways to handle data storage — a B-Tree (a generalization of Binary Search Tree) and a Log-Structured Merge Tree, both with having pros and cons.
Here are a few short articles to get a grasp of the trade-offs between the two:
1. https://blog.yugabyte.com/a-busy-developers-guide-to-database-storage-engines-the-basics/
2. https://blog.yugabyte.com/a-busy-developers-guide-to-database-storage-engines-advanced-topics
3. https://rkenmi.com/posts/b-trees-vs-lsm-trees
4. https://tikv.org/deep-dive/key-value-engine/b-tree-vs-lsm/
#systemdesign #dbs #databases #design #btree #lsmtree #storage #performance #sql #nosql #engine
đź’» The byte is not enough
Have you ever concern yourself with the question of how modern database systems' internal storage works? Well, there are two popular ways to handle data storage — a B-Tree (a generalization of Binary Search Tree) and a Log-Structured Merge Tree, both with having pros and cons.
Here are a few short articles to get a grasp of the trade-offs between the two:
1. https://blog.yugabyte.com/a-busy-developers-guide-to-database-storage-engines-the-basics/
2. https://blog.yugabyte.com/a-busy-developers-guide-to-database-storage-engines-advanced-topics
3. https://rkenmi.com/posts/b-trees-vs-lsm-trees
4. https://tikv.org/deep-dive/key-value-engine/b-tree-vs-lsm/
#systemdesign #dbs #databases #design #btree #lsmtree #storage #performance #sql #nosql #engine
đź’» The byte is not enough
zab.totally-ordered-broadcast-protocol.2008.pdf
264.7 KB
🧑‍🔬 Architecture of ZAB – ZooKeeper Atomic Broadcast protocol
The ZAB protocol ensures that the Zookeeper replication is done in order and is also responsible for the election of leader nodes and the restoration of any failed nodes. In a Zookeeper ecosystem, the leader node is the heart of everything; every cluster has one leader node and the rest of the nodes are followers. All incoming client requests and state changes are received at first by the leader with responsibility to replicate it across all its followers (and itself). All incoming read requests are also load balanced by the leader within itself and its followers.
Original ZAB paper:
https://marcoserafini.github.io/papers/zab.pdf
Implementation details of ZAB:
http://www.tcs.hut.fi/Studies/T-79.5001/reports/2012-deSouzaMedeiros.pdf
#systemdesign #zab #consensus #total #order #broadcast #design #quorum #performance #faulttolerance #yahoo #paper #research #scalability
đź’» The byte is not enough
The ZAB protocol ensures that the Zookeeper replication is done in order and is also responsible for the election of leader nodes and the restoration of any failed nodes. In a Zookeeper ecosystem, the leader node is the heart of everything; every cluster has one leader node and the rest of the nodes are followers. All incoming client requests and state changes are received at first by the leader with responsibility to replicate it across all its followers (and itself). All incoming read requests are also load balanced by the leader within itself and its followers.
Original ZAB paper:
https://marcoserafini.github.io/papers/zab.pdf
Implementation details of ZAB:
http://www.tcs.hut.fi/Studies/T-79.5001/reports/2012-deSouzaMedeiros.pdf
#systemdesign #zab #consensus #total #order #broadcast #design #quorum #performance #faulttolerance #yahoo #paper #research #scalability
đź’» The byte is not enough
Building Facebook’s service encryption infrastructure
What does it take to incorporate security protocols into a system of millions of services running on thousands of machines across the world? You could use some well-known technologies as Kerberos, but will soon realize it does not work on a big scale. You would then probably stick to the idea of building a custom in-house solution, which is exactly what Facebook did to suit their needs in security and operability without compromising performance.
https://engineering.fb.com/2019/05/29/security/service-encryption/
#systemdesign #security #facebook #kerberos #tls #encryption #traffic #networking #dns #performance #faulttolerance #scalability #infrastructure
đź’» The byte is not enough
What does it take to incorporate security protocols into a system of millions of services running on thousands of machines across the world? You could use some well-known technologies as Kerberos, but will soon realize it does not work on a big scale. You would then probably stick to the idea of building a custom in-house solution, which is exactly what Facebook did to suit their needs in security and operability without compromising performance.
https://engineering.fb.com/2019/05/29/security/service-encryption/
#systemdesign #security #facebook #kerberos #tls #encryption #traffic #networking #dns #performance #faulttolerance #scalability #infrastructure
đź’» The byte is not enough
Engineering at Meta
Building Facebook’s service encryption infrastructure
We run one of the largest microservices deployments in the world, with thousands of services that perform billions of requests per second. Keeping information secure as these services communicate g…
🧑‍🔬 CORFU: A Distributed Shared Log
Despite almost forty years of research into replicated storage schemes, the only approach so far to scale up capacity and throughput has been to shard data and trade consistency for performance. The CORFU system breaks this seeming tradeoff by organizing a cluster of drives as a single, shared log. CORFU offers a single-copy semantics at cluster-scale speeds, providing a scalable source of atomicity and durability for distributed systems.
https://www.youtube.com/watch?v=GmVQVT9aZfU
Original paper:
https://www.cs.utexas.edu/~lorenzo/corsi/cs380d/papers/a10-balakrishnan.pdf
#systemdesign #corfu #microsoft #design #reliability #algorithms #distributed #log #replication #consensus #performance #scalability #faulttolerance #paper #research #infrastructure
đź’» The byte is not enough
Despite almost forty years of research into replicated storage schemes, the only approach so far to scale up capacity and throughput has been to shard data and trade consistency for performance. The CORFU system breaks this seeming tradeoff by organizing a cluster of drives as a single, shared log. CORFU offers a single-copy semantics at cluster-scale speeds, providing a scalable source of atomicity and durability for distributed systems.
https://www.youtube.com/watch?v=GmVQVT9aZfU
Original paper:
https://www.cs.utexas.edu/~lorenzo/corsi/cs380d/papers/a10-balakrishnan.pdf
#systemdesign #corfu #microsoft #design #reliability #algorithms #distributed #log #replication #consensus #performance #scalability #faulttolerance #paper #research #infrastructure
đź’» The byte is not enough
YouTube
Michael Wei - Corfu: A Cloud-Scale Consistency Platform
Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.
🧑‍🔬 Delos: Simple, flexible storage for the Facebook control plane
Facebook Delos — is a fundamentally new architecture for building replicated storage systems. Its modular, layered design provides flexibility and simplicity without sacrificing performance or reliability. Delos enables fast time-to-deployment for new storage systems — Facebook deployed an initial version in production within eight months — as well as safe, rapid evolution. Facebook swapped in a new ordering mechanism to obtain 10x lower latency without any service downtime.
https://www.youtube.com/watch?v=wd-GC_XhA2g&t=314s
Blog post presentation:
https://engineering.fb.com/2019/06/06/data-center-engineering/delos/
Original paper:
https://www.usenix.org/system/files/osdi20-balakrishnan.pdf
#systemdesign #delos #facebook #distributed #shared #log #algorithms #design #consensus #performance #scalability #faulttolerance #paper #research #infrastructure #storage #controlplane #corfu #paxos #zookeeper #zab
đź’» The byte is not enough
Facebook Delos — is a fundamentally new architecture for building replicated storage systems. Its modular, layered design provides flexibility and simplicity without sacrificing performance or reliability. Delos enables fast time-to-deployment for new storage systems — Facebook deployed an initial version in production within eight months — as well as safe, rapid evolution. Facebook swapped in a new ordering mechanism to obtain 10x lower latency without any service downtime.
https://www.youtube.com/watch?v=wd-GC_XhA2g&t=314s
Blog post presentation:
https://engineering.fb.com/2019/06/06/data-center-engineering/delos/
Original paper:
https://www.usenix.org/system/files/osdi20-balakrishnan.pdf
#systemdesign #delos #facebook #distributed #shared #log #algorithms #design #consensus #performance #scalability #faulttolerance #paper #research #infrastructure #storage #controlplane #corfu #paxos #zookeeper #zab
đź’» The byte is not enough
YouTube
OSDI '20 - Virtual Consensus in Delos
Virtual Consensus in Delos
Mahesh Balakrishnan, Jason Flinn, Chen Shen, Mihir Dharamshi, Ahmed Jafri, Xiao Shi, Santosh Ghosh, Hazem Hassan, Aaryaman Sagar, Rhed Shi, Jingming Liu, Filip Gruszczynski, Xianan Zhang, Huy Hoang, Ahmed Yossef, Francois Richard…
Mahesh Balakrishnan, Jason Flinn, Chen Shen, Mihir Dharamshi, Ahmed Jafri, Xiao Shi, Santosh Ghosh, Hazem Hassan, Aaryaman Sagar, Rhed Shi, Jingming Liu, Filip Gruszczynski, Xianan Zhang, Huy Hoang, Ahmed Yossef, Francois Richard…
Building a real-time user action counting system for ads
A quick read on how Pinterest built an ad tracking platform called Aperture to count the number of ad interactions in the past one, three, or 30 days.
https://medium.com/pinterest-engineering/building-a-real-time-user-action-counting-system-for-ads-88a60d9c9a
#systemdesign #pinterest #aperture #kafka #distributed #counter #infrastructure #tech #blog #adtech #ads #rocksdb
đź’» The byte is not enough
A quick read on how Pinterest built an ad tracking platform called Aperture to count the number of ad interactions in the past one, three, or 30 days.
https://medium.com/pinterest-engineering/building-a-real-time-user-action-counting-system-for-ads-88a60d9c9a
#systemdesign #pinterest #aperture #kafka #distributed #counter #infrastructure #tech #blog #adtech #ads #rocksdb
đź’» The byte is not enough
Medium
Building a real-time user action counting system for ads
Del Bao, Software engineer, Ads Infrastructure
Guodong Han and Jian Fang, Software engineers, Serving System
Guodong Han and Jian Fang, Software engineers, Serving System
Geofencing: An Overview
An interesting benchmarked overview of some widely used geofencing algorithms such as RTree and QuadTree in response to a very controversial post at Uber's engineering blog.
https://medium.com/@buckhx/unwinding-uber-s-most-efficient-service-406413c5871d
Original Uber post:
https://eng.uber.com/go-geofence-highest-query-per-second-service/
#algorithms #geofencing #uber #quadtree #rtree #kdtree #s2 #maps #location #performance #golang
đź’» The byte is not enough
An interesting benchmarked overview of some widely used geofencing algorithms such as RTree and QuadTree in response to a very controversial post at Uber's engineering blog.
https://medium.com/@buckhx/unwinding-uber-s-most-efficient-service-406413c5871d
Original Uber post:
https://eng.uber.com/go-geofence-highest-query-per-second-service/
#algorithms #geofencing #uber #quadtree #rtree #kdtree #s2 #maps #location #performance #golang
đź’» The byte is not enough
Medium
Unwinding Uber’s Most Efficient Service
A few weeks ago, Uber posted an article detailing how they built their “highest query per second service using Go”. The article is fairly…
Shallow Mirror
Example usage of Kafka's MirrorMaker at scale (at Pinterest) and potential performance issues it may bring into the system (and how they could be fixed effectively).
https://medium.com/pinterest-engineering/shallow-mirror-f543b14bb25
#systemdesign #pinterest #kafka #queue #shared #log #mirrormaker #infrastructure #commit #performance #scalability #reliability
đź’» The byte is not enough
Example usage of Kafka's MirrorMaker at scale (at Pinterest) and potential performance issues it may bring into the system (and how they could be fixed effectively).
https://medium.com/pinterest-engineering/shallow-mirror-f543b14bb25
#systemdesign #pinterest #kafka #queue #shared #log #mirrormaker #infrastructure #commit #performance #scalability #reliability
đź’» The byte is not enough
Medium
Shallow Mirror
Enhancement to Kafka MirrorMaker to reduce CPU/memory pressure
Google’s S2, geometry on the sphere, cells and Hilbert curve
An informative introduction on how Google leverages geometry concepts like Hilbert curves in their famous S2 library that offers spatial indexing capabilities tuned for performance and correctness. Python code examples included.
https://blog.christianperone.com/2015/08/googles-s2-geometry-on-the-sphere-cells-and-hilbert-curve/
Hilbert curves introduction:
https://datagenetics.com/blog/march22013/index.html
#algorithms #math #geometry #python #performance #geo #fencing #spatial #index #s2 #rtree #quadtree #google
đź’» The byte is not enough
An informative introduction on how Google leverages geometry concepts like Hilbert curves in their famous S2 library that offers spatial indexing capabilities tuned for performance and correctness. Python code examples included.
https://blog.christianperone.com/2015/08/googles-s2-geometry-on-the-sphere-cells-and-hilbert-curve/
Hilbert curves introduction:
https://datagenetics.com/blog/march22013/index.html
#algorithms #math #geometry #python #performance #geo #fencing #spatial #index #s2 #rtree #quadtree #google
đź’» The byte is not enough
Christianperone
Google’s S2, geometry on the sphere, cells and Hilbert curve | Terra Incognita
Update - 05 Dec 2017: Google just announced that it will be commited to the development of a new released version of the S2 library, amazing news, repository can be found here. Google's S2 library is a real treasure, not only due to its capabilities for spatial…
Fighting spam with Guardian, a real-time analytics and rules engine
Another great story of a practical approach to fight spam at big scale that is being applied at Pinterest. Though this isn't a story of inventing some tremendous and big all-in-one solution, it still provides a nice perspective on how to grow from a small and clumsy solution to a big and robust killer feature.
https://medium.com/pinterest-engineering/fighting-spam-with-guardian-a-real-time-analytics-and-rules-engine-938e7e61fa27
#systemdesign #infrastructure #pinterest #design #scalability #performance #bigdata #hive #kafka #presto #guardian #spam #realtime #analytics #safety #security
đź’» The byte is not enough
Another great story of a practical approach to fight spam at big scale that is being applied at Pinterest. Though this isn't a story of inventing some tremendous and big all-in-one solution, it still provides a nice perspective on how to grow from a small and clumsy solution to a big and robust killer feature.
https://medium.com/pinterest-engineering/fighting-spam-with-guardian-a-real-time-analytics-and-rules-engine-938e7e61fa27
#systemdesign #infrastructure #pinterest #design #scalability #performance #bigdata #hive #kafka #presto #guardian #spam #realtime #analytics #safety #security
đź’» The byte is not enough
Medium
Fighting spam with Guardian, a real-time analytics and rules engine
Hongkai Pan | Software Engineer, Trust & Safety
(Strong) Eventual Consistency and Conflict-free Replicated Data Types (CRDT)
Conflict-free Replicated Data Types (CRDTs) are an increasingly popular family of algorithms for optimistic replication. They allow data to be concurrently updated on several replicas, even while those replicas are offline, and provide a robust way of merging those updates back into a consistent state. CRDTs are used in geo-replicated databases, multi-user collaboration software, distributed processing frameworks, and various other systems.
However, while the basic principles of CRDTs are now quite well known, many challenging problems are lurking below the surface. It turns out that CRDTs are easy to implement badly. Many published algorithms have anomalies that cause them to behave strangely in some situations. Simple implementations often have terrible performance, and making the performance good is challenging.
Martin Kleppmann's talk on CRDT at Hydra Conf:
https://www.youtube.com/watch?v=PMVBuMK_pJY
Sean Cribbs discusses Convergent Replicated Data Types, data structures that tolerate eventual consistency:
https://www.infoq.com/presentations/CRDT/
Marc Shapiro presentation on SEC and CRTD at Microsoft Research:
https://www.youtube.com/watch?v=oyUHd894w18
#systemdesign #algorithms #sec #strong #consistency #crdt #replication #consensus #conflicts #resolution #riak #math #distributed
đź’» The byte is not enough
Conflict-free Replicated Data Types (CRDTs) are an increasingly popular family of algorithms for optimistic replication. They allow data to be concurrently updated on several replicas, even while those replicas are offline, and provide a robust way of merging those updates back into a consistent state. CRDTs are used in geo-replicated databases, multi-user collaboration software, distributed processing frameworks, and various other systems.
However, while the basic principles of CRDTs are now quite well known, many challenging problems are lurking below the surface. It turns out that CRDTs are easy to implement badly. Many published algorithms have anomalies that cause them to behave strangely in some situations. Simple implementations often have terrible performance, and making the performance good is challenging.
Martin Kleppmann's talk on CRDT at Hydra Conf:
https://www.youtube.com/watch?v=PMVBuMK_pJY
Sean Cribbs discusses Convergent Replicated Data Types, data structures that tolerate eventual consistency:
https://www.infoq.com/presentations/CRDT/
Marc Shapiro presentation on SEC and CRTD at Microsoft Research:
https://www.youtube.com/watch?v=oyUHd894w18
#systemdesign #algorithms #sec #strong #consistency #crdt #replication #consensus #conflicts #resolution #riak #math #distributed
đź’» The byte is not enough
YouTube
Martin Kleppmann — CRDTs: The hard parts
About Hydra conference: https://jrg.su/6Cf8RP
— Hydra 2022 — June 2-3
Info and tickets: https://bit.ly/3ni5Hem
— —
Conflict-free Replicated Data Types (CRDTs) are an increasingly popular family of algorithms for optimistic replication. They allow data to…
— Hydra 2022 — June 2-3
Info and tickets: https://bit.ly/3ni5Hem
— —
Conflict-free Replicated Data Types (CRDTs) are an increasingly popular family of algorithms for optimistic replication. They allow data to…
Building scalable near-real time indexing on HBase
An arguably simple and elegant approach to enabling near real-time indexing in HBase that took place at Pinterest.
https://medium.com/pinterest-engineering/building-scalable-near-real-time-indexing-on-hbase-7b5eeb411888
#systemdesign #pinterest #hbasee #hdfs #indexing #search #faulttolerance #scalability #performance #kafka #cache #consistency #infrastructure #design
đź’» The byte is not enough
An arguably simple and elegant approach to enabling near real-time indexing in HBase that took place at Pinterest.
https://medium.com/pinterest-engineering/building-scalable-near-real-time-indexing-on-hbase-7b5eeb411888
#systemdesign #pinterest #hbasee #hdfs #indexing #search #faulttolerance #scalability #performance #kafka #cache #consistency #infrastructure #design
đź’» The byte is not enough
Medium
Building scalable near-real time indexing on HBase
Ankita Wagh | Software Engineer, Storage and Caching
đź•° Blast from the past: Please stop calling databases CP or AP
A post from 2015 where Martin Kleppmann argues about usage of CAP-theorem in context of modern distributed systems. Martin does a great job demonstrating that people tend to attribute terms like consistency and availability with the meaning that differ from the one given in the original Brewer's CAP Theorem paper, and as thus demonstrates that almost no modern distributed system could be characterized with "A" or "C" property.
https://martin.kleppmann.com/2015/05/11/please-stop-calling-databases-cp-or-ap.html
#systemdesign #theory #research #cap #brewer #kleppmann #consistency #availability #partitiontolerance #zab #consensus #linearizability #zookeeper #riak #dynamodb #cassandra #voldemort #cp #aperture
đź’» The byte is not enough
A post from 2015 where Martin Kleppmann argues about usage of CAP-theorem in context of modern distributed systems. Martin does a great job demonstrating that people tend to attribute terms like consistency and availability with the meaning that differ from the one given in the original Brewer's CAP Theorem paper, and as thus demonstrates that almost no modern distributed system could be characterized with "A" or "C" property.
https://martin.kleppmann.com/2015/05/11/please-stop-calling-databases-cp-or-ap.html
#systemdesign #theory #research #cap #brewer #kleppmann #consistency #availability #partitiontolerance #zab #consensus #linearizability #zookeeper #riak #dynamodb #cassandra #voldemort #cp #aperture
đź’» The byte is not enough