🧑‍🔬 CORFU: A Distributed Shared Log
Despite almost forty years of research into replicated storage schemes, the only approach so far to scale up capacity and throughput has been to shard data and trade consistency for performance. The CORFU system breaks this seeming tradeoff by organizing a cluster of drives as a single, shared log. CORFU offers a single-copy semantics at cluster-scale speeds, providing a scalable source of atomicity and durability for distributed systems.
https://www.youtube.com/watch?v=GmVQVT9aZfU
Original paper:
https://www.cs.utexas.edu/~lorenzo/corsi/cs380d/papers/a10-balakrishnan.pdf
#systemdesign #corfu #microsoft #design #reliability #algorithms #distributed #log #replication #consensus #performance #scalability #faulttolerance #paper #research #infrastructure
đź’» The byte is not enough
Despite almost forty years of research into replicated storage schemes, the only approach so far to scale up capacity and throughput has been to shard data and trade consistency for performance. The CORFU system breaks this seeming tradeoff by organizing a cluster of drives as a single, shared log. CORFU offers a single-copy semantics at cluster-scale speeds, providing a scalable source of atomicity and durability for distributed systems.
https://www.youtube.com/watch?v=GmVQVT9aZfU
Original paper:
https://www.cs.utexas.edu/~lorenzo/corsi/cs380d/papers/a10-balakrishnan.pdf
#systemdesign #corfu #microsoft #design #reliability #algorithms #distributed #log #replication #consensus #performance #scalability #faulttolerance #paper #research #infrastructure
đź’» The byte is not enough
YouTube
Michael Wei - Corfu: A Cloud-Scale Consistency Platform
Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.
🧑‍🔬 Delos: Simple, flexible storage for the Facebook control plane
Facebook Delos — is a fundamentally new architecture for building replicated storage systems. Its modular, layered design provides flexibility and simplicity without sacrificing performance or reliability. Delos enables fast time-to-deployment for new storage systems — Facebook deployed an initial version in production within eight months — as well as safe, rapid evolution. Facebook swapped in a new ordering mechanism to obtain 10x lower latency without any service downtime.
https://www.youtube.com/watch?v=wd-GC_XhA2g&t=314s
Blog post presentation:
https://engineering.fb.com/2019/06/06/data-center-engineering/delos/
Original paper:
https://www.usenix.org/system/files/osdi20-balakrishnan.pdf
#systemdesign #delos #facebook #distributed #shared #log #algorithms #design #consensus #performance #scalability #faulttolerance #paper #research #infrastructure #storage #controlplane #corfu #paxos #zookeeper #zab
đź’» The byte is not enough
Facebook Delos — is a fundamentally new architecture for building replicated storage systems. Its modular, layered design provides flexibility and simplicity without sacrificing performance or reliability. Delos enables fast time-to-deployment for new storage systems — Facebook deployed an initial version in production within eight months — as well as safe, rapid evolution. Facebook swapped in a new ordering mechanism to obtain 10x lower latency without any service downtime.
https://www.youtube.com/watch?v=wd-GC_XhA2g&t=314s
Blog post presentation:
https://engineering.fb.com/2019/06/06/data-center-engineering/delos/
Original paper:
https://www.usenix.org/system/files/osdi20-balakrishnan.pdf
#systemdesign #delos #facebook #distributed #shared #log #algorithms #design #consensus #performance #scalability #faulttolerance #paper #research #infrastructure #storage #controlplane #corfu #paxos #zookeeper #zab
đź’» The byte is not enough
YouTube
OSDI '20 - Virtual Consensus in Delos
Virtual Consensus in Delos
Mahesh Balakrishnan, Jason Flinn, Chen Shen, Mihir Dharamshi, Ahmed Jafri, Xiao Shi, Santosh Ghosh, Hazem Hassan, Aaryaman Sagar, Rhed Shi, Jingming Liu, Filip Gruszczynski, Xianan Zhang, Huy Hoang, Ahmed Yossef, Francois Richard…
Mahesh Balakrishnan, Jason Flinn, Chen Shen, Mihir Dharamshi, Ahmed Jafri, Xiao Shi, Santosh Ghosh, Hazem Hassan, Aaryaman Sagar, Rhed Shi, Jingming Liu, Filip Gruszczynski, Xianan Zhang, Huy Hoang, Ahmed Yossef, Francois Richard…
Building a real-time user action counting system for ads
A quick read on how Pinterest built an ad tracking platform called Aperture to count the number of ad interactions in the past one, three, or 30 days.
https://medium.com/pinterest-engineering/building-a-real-time-user-action-counting-system-for-ads-88a60d9c9a
#systemdesign #pinterest #aperture #kafka #distributed #counter #infrastructure #tech #blog #adtech #ads #rocksdb
đź’» The byte is not enough
A quick read on how Pinterest built an ad tracking platform called Aperture to count the number of ad interactions in the past one, three, or 30 days.
https://medium.com/pinterest-engineering/building-a-real-time-user-action-counting-system-for-ads-88a60d9c9a
#systemdesign #pinterest #aperture #kafka #distributed #counter #infrastructure #tech #blog #adtech #ads #rocksdb
đź’» The byte is not enough
Medium
Building a real-time user action counting system for ads
Del Bao, Software engineer, Ads Infrastructure
Guodong Han and Jian Fang, Software engineers, Serving System
Guodong Han and Jian Fang, Software engineers, Serving System
Geofencing: An Overview
An interesting benchmarked overview of some widely used geofencing algorithms such as RTree and QuadTree in response to a very controversial post at Uber's engineering blog.
https://medium.com/@buckhx/unwinding-uber-s-most-efficient-service-406413c5871d
Original Uber post:
https://eng.uber.com/go-geofence-highest-query-per-second-service/
#algorithms #geofencing #uber #quadtree #rtree #kdtree #s2 #maps #location #performance #golang
đź’» The byte is not enough
An interesting benchmarked overview of some widely used geofencing algorithms such as RTree and QuadTree in response to a very controversial post at Uber's engineering blog.
https://medium.com/@buckhx/unwinding-uber-s-most-efficient-service-406413c5871d
Original Uber post:
https://eng.uber.com/go-geofence-highest-query-per-second-service/
#algorithms #geofencing #uber #quadtree #rtree #kdtree #s2 #maps #location #performance #golang
đź’» The byte is not enough
Medium
Unwinding Uber’s Most Efficient Service
A few weeks ago, Uber posted an article detailing how they built their “highest query per second service using Go”. The article is fairly…
Shallow Mirror
Example usage of Kafka's MirrorMaker at scale (at Pinterest) and potential performance issues it may bring into the system (and how they could be fixed effectively).
https://medium.com/pinterest-engineering/shallow-mirror-f543b14bb25
#systemdesign #pinterest #kafka #queue #shared #log #mirrormaker #infrastructure #commit #performance #scalability #reliability
đź’» The byte is not enough
Example usage of Kafka's MirrorMaker at scale (at Pinterest) and potential performance issues it may bring into the system (and how they could be fixed effectively).
https://medium.com/pinterest-engineering/shallow-mirror-f543b14bb25
#systemdesign #pinterest #kafka #queue #shared #log #mirrormaker #infrastructure #commit #performance #scalability #reliability
đź’» The byte is not enough
Medium
Shallow Mirror
Enhancement to Kafka MirrorMaker to reduce CPU/memory pressure
Google’s S2, geometry on the sphere, cells and Hilbert curve
An informative introduction on how Google leverages geometry concepts like Hilbert curves in their famous S2 library that offers spatial indexing capabilities tuned for performance and correctness. Python code examples included.
https://blog.christianperone.com/2015/08/googles-s2-geometry-on-the-sphere-cells-and-hilbert-curve/
Hilbert curves introduction:
https://datagenetics.com/blog/march22013/index.html
#algorithms #math #geometry #python #performance #geo #fencing #spatial #index #s2 #rtree #quadtree #google
đź’» The byte is not enough
An informative introduction on how Google leverages geometry concepts like Hilbert curves in their famous S2 library that offers spatial indexing capabilities tuned for performance and correctness. Python code examples included.
https://blog.christianperone.com/2015/08/googles-s2-geometry-on-the-sphere-cells-and-hilbert-curve/
Hilbert curves introduction:
https://datagenetics.com/blog/march22013/index.html
#algorithms #math #geometry #python #performance #geo #fencing #spatial #index #s2 #rtree #quadtree #google
đź’» The byte is not enough
Christianperone
Google’s S2, geometry on the sphere, cells and Hilbert curve | Terra Incognita
Update - 05 Dec 2017: Google just announced that it will be commited to the development of a new released version of the S2 library, amazing news, repository can be found here. Google's S2 library is a real treasure, not only due to its capabilities for spatial…
Fighting spam with Guardian, a real-time analytics and rules engine
Another great story of a practical approach to fight spam at big scale that is being applied at Pinterest. Though this isn't a story of inventing some tremendous and big all-in-one solution, it still provides a nice perspective on how to grow from a small and clumsy solution to a big and robust killer feature.
https://medium.com/pinterest-engineering/fighting-spam-with-guardian-a-real-time-analytics-and-rules-engine-938e7e61fa27
#systemdesign #infrastructure #pinterest #design #scalability #performance #bigdata #hive #kafka #presto #guardian #spam #realtime #analytics #safety #security
đź’» The byte is not enough
Another great story of a practical approach to fight spam at big scale that is being applied at Pinterest. Though this isn't a story of inventing some tremendous and big all-in-one solution, it still provides a nice perspective on how to grow from a small and clumsy solution to a big and robust killer feature.
https://medium.com/pinterest-engineering/fighting-spam-with-guardian-a-real-time-analytics-and-rules-engine-938e7e61fa27
#systemdesign #infrastructure #pinterest #design #scalability #performance #bigdata #hive #kafka #presto #guardian #spam #realtime #analytics #safety #security
đź’» The byte is not enough
Medium
Fighting spam with Guardian, a real-time analytics and rules engine
Hongkai Pan | Software Engineer, Trust & Safety
(Strong) Eventual Consistency and Conflict-free Replicated Data Types (CRDT)
Conflict-free Replicated Data Types (CRDTs) are an increasingly popular family of algorithms for optimistic replication. They allow data to be concurrently updated on several replicas, even while those replicas are offline, and provide a robust way of merging those updates back into a consistent state. CRDTs are used in geo-replicated databases, multi-user collaboration software, distributed processing frameworks, and various other systems.
However, while the basic principles of CRDTs are now quite well known, many challenging problems are lurking below the surface. It turns out that CRDTs are easy to implement badly. Many published algorithms have anomalies that cause them to behave strangely in some situations. Simple implementations often have terrible performance, and making the performance good is challenging.
Martin Kleppmann's talk on CRDT at Hydra Conf:
https://www.youtube.com/watch?v=PMVBuMK_pJY
Sean Cribbs discusses Convergent Replicated Data Types, data structures that tolerate eventual consistency:
https://www.infoq.com/presentations/CRDT/
Marc Shapiro presentation on SEC and CRTD at Microsoft Research:
https://www.youtube.com/watch?v=oyUHd894w18
#systemdesign #algorithms #sec #strong #consistency #crdt #replication #consensus #conflicts #resolution #riak #math #distributed
đź’» The byte is not enough
Conflict-free Replicated Data Types (CRDTs) are an increasingly popular family of algorithms for optimistic replication. They allow data to be concurrently updated on several replicas, even while those replicas are offline, and provide a robust way of merging those updates back into a consistent state. CRDTs are used in geo-replicated databases, multi-user collaboration software, distributed processing frameworks, and various other systems.
However, while the basic principles of CRDTs are now quite well known, many challenging problems are lurking below the surface. It turns out that CRDTs are easy to implement badly. Many published algorithms have anomalies that cause them to behave strangely in some situations. Simple implementations often have terrible performance, and making the performance good is challenging.
Martin Kleppmann's talk on CRDT at Hydra Conf:
https://www.youtube.com/watch?v=PMVBuMK_pJY
Sean Cribbs discusses Convergent Replicated Data Types, data structures that tolerate eventual consistency:
https://www.infoq.com/presentations/CRDT/
Marc Shapiro presentation on SEC and CRTD at Microsoft Research:
https://www.youtube.com/watch?v=oyUHd894w18
#systemdesign #algorithms #sec #strong #consistency #crdt #replication #consensus #conflicts #resolution #riak #math #distributed
đź’» The byte is not enough
YouTube
Martin Kleppmann — CRDTs: The hard parts
About Hydra conference: https://jrg.su/6Cf8RP
— Hydra 2022 — June 2-3
Info and tickets: https://bit.ly/3ni5Hem
— —
Conflict-free Replicated Data Types (CRDTs) are an increasingly popular family of algorithms for optimistic replication. They allow data to…
— Hydra 2022 — June 2-3
Info and tickets: https://bit.ly/3ni5Hem
— —
Conflict-free Replicated Data Types (CRDTs) are an increasingly popular family of algorithms for optimistic replication. They allow data to…
Building scalable near-real time indexing on HBase
An arguably simple and elegant approach to enabling near real-time indexing in HBase that took place at Pinterest.
https://medium.com/pinterest-engineering/building-scalable-near-real-time-indexing-on-hbase-7b5eeb411888
#systemdesign #pinterest #hbasee #hdfs #indexing #search #faulttolerance #scalability #performance #kafka #cache #consistency #infrastructure #design
đź’» The byte is not enough
An arguably simple and elegant approach to enabling near real-time indexing in HBase that took place at Pinterest.
https://medium.com/pinterest-engineering/building-scalable-near-real-time-indexing-on-hbase-7b5eeb411888
#systemdesign #pinterest #hbasee #hdfs #indexing #search #faulttolerance #scalability #performance #kafka #cache #consistency #infrastructure #design
đź’» The byte is not enough
Medium
Building scalable near-real time indexing on HBase
Ankita Wagh | Software Engineer, Storage and Caching
đź•° Blast from the past: Please stop calling databases CP or AP
A post from 2015 where Martin Kleppmann argues about usage of CAP-theorem in context of modern distributed systems. Martin does a great job demonstrating that people tend to attribute terms like consistency and availability with the meaning that differ from the one given in the original Brewer's CAP Theorem paper, and as thus demonstrates that almost no modern distributed system could be characterized with "A" or "C" property.
https://martin.kleppmann.com/2015/05/11/please-stop-calling-databases-cp-or-ap.html
#systemdesign #theory #research #cap #brewer #kleppmann #consistency #availability #partitiontolerance #zab #consensus #linearizability #zookeeper #riak #dynamodb #cassandra #voldemort #cp #aperture
đź’» The byte is not enough
A post from 2015 where Martin Kleppmann argues about usage of CAP-theorem in context of modern distributed systems. Martin does a great job demonstrating that people tend to attribute terms like consistency and availability with the meaning that differ from the one given in the original Brewer's CAP Theorem paper, and as thus demonstrates that almost no modern distributed system could be characterized with "A" or "C" property.
https://martin.kleppmann.com/2015/05/11/please-stop-calling-databases-cp-or-ap.html
#systemdesign #theory #research #cap #brewer #kleppmann #consistency #availability #partitiontolerance #zab #consensus #linearizability #zookeeper #riak #dynamodb #cassandra #voldemort #cp #aperture
đź’» The byte is not enough
Understanding Lamport Timestamps with Python’s multiprocessing library
A brief introduction into Lamport Timestamps and its implementation known as Vector Clocks that helps to distinguish casual ordering between events. Python code examples included.
https://towardsdatascience.com/understanding-lamport-timestamps-with-pythons-multiprocessing-library-12a6427881c6
#algorithms #systemdesign #distributed #ordering #clocks #vector #logical #python #research #paper #lamport #casualordering #synchronization #partial #concurrent #timestamps
đź’» The byte is not enough
A brief introduction into Lamport Timestamps and its implementation known as Vector Clocks that helps to distinguish casual ordering between events. Python code examples included.
https://towardsdatascience.com/understanding-lamport-timestamps-with-pythons-multiprocessing-library-12a6427881c6
#algorithms #systemdesign #distributed #ordering #clocks #vector #logical #python #research #paper #lamport #casualordering #synchronization #partial #concurrent #timestamps
đź’» The byte is not enough
Medium
Understanding Lamport Timestamps with Python’s multiprocessing library
Everyone who has been working with distributed systems or logs from such a systems, has directly or indirectly encountered Lamport…
System Design Interview: The Bare Minimum You Need to Know about Load Balancers
Small introduction to basic things you should know about load balancers before you go for your system design interview.
https://medium.com/swlh/system-design-interview-the-bare-minimum-you-need-to-know-about-load-balancers-fc0cbe1ac276
#systemdesign #sdi #loadbalancers #performance #design #infrastructure #spof #ssl #redundancy
đź’» The byte is not enough
Small introduction to basic things you should know about load balancers before you go for your system design interview.
https://medium.com/swlh/system-design-interview-the-bare-minimum-you-need-to-know-about-load-balancers-fc0cbe1ac276
#systemdesign #sdi #loadbalancers #performance #design #infrastructure #spof #ssl #redundancy
đź’» The byte is not enough
Medium
System Design Interview: The Bare Minimum You Need to Know about Load Balancers
I still remember the early day when I failed a dream job interview:
6 Things You Need to Know About Kafka Before Using it in a System Design Interview
In a typical system design round, there's a chance you may want to leverage Kafka capabilities in your designated system. But what if an interviewer start asking you questions about Kafka's inner mechanics? Will you be able to answer those questions? Here is a list of 6 most used/interesting aspects of Kafka, knowing which will greatly improve your overall system design skills.
https://levelup.gitconnected.com/6-things-you-need-to-know-about-kafka-before-using-it-in-a-system-design-interview-1fc31451732c
#systemdesign #kafka #distributed #queue #log #infrastructure #design #shared #commit #consumer #producer #faulttolerance #reliability #scalability #performance #broker #zookeeper
đź’» The byte is not enough
In a typical system design round, there's a chance you may want to leverage Kafka capabilities in your designated system. But what if an interviewer start asking you questions about Kafka's inner mechanics? Will you be able to answer those questions? Here is a list of 6 most used/interesting aspects of Kafka, knowing which will greatly improve your overall system design skills.
https://levelup.gitconnected.com/6-things-you-need-to-know-about-kafka-before-using-it-in-a-system-design-interview-1fc31451732c
#systemdesign #kafka #distributed #queue #log #infrastructure #design #shared #commit #consumer #producer #faulttolerance #reliability #scalability #performance #broker #zookeeper
đź’» The byte is not enough
Medium
6 Things You Need to Know About Kafka Before Using it in a System Design Interview
I quite often ran into system design discussions in which people seemed to have the tendency to use Kafka as a magic box, as if drawing…
How Google Spanner Assigns Commit Timestamps — The Secret Sauce of Its Strong Consistency
Spanner is Google’s global scale and synchronously replicated relational database. Its research paper was first published in 2012. Spanner received much acclaim because it was the first system to distribute data at global scale and support strongly consistent distributed transactions. That seems to break the CAP theorem. But in fact, Spanner chooses consistency over availability in face of network partition. It’s just that Google’s data center infrastructure is so reliable that external users typically don’t worry about its outages. Google now offers Spanner as a service through its cloud platform.
https://levelup.gitconnected.com/how-google-spanner-assigns-commit-timestamps-the-secret-sauce-of-its-strong-consistency-8bc143614f26
#systemdesign #google #spanner #transactions #commit #distributed #database #infrastructure #scalability #consistency #quorum #performance #replication #truetime #api #clocks #synchronization #atomicity
đź’» The byte is not enough
Spanner is Google’s global scale and synchronously replicated relational database. Its research paper was first published in 2012. Spanner received much acclaim because it was the first system to distribute data at global scale and support strongly consistent distributed transactions. That seems to break the CAP theorem. But in fact, Spanner chooses consistency over availability in face of network partition. It’s just that Google’s data center infrastructure is so reliable that external users typically don’t worry about its outages. Google now offers Spanner as a service through its cloud platform.
https://levelup.gitconnected.com/how-google-spanner-assigns-commit-timestamps-the-secret-sauce-of-its-strong-consistency-8bc143614f26
#systemdesign #google #spanner #transactions #commit #distributed #database #infrastructure #scalability #consistency #quorum #performance #replication #truetime #api #clocks #synchronization #atomicity
đź’» The byte is not enough
Medium
How Google Spanner Assigns Commit Timestamps — The Secret Sauce of Its Strong Consistency
Spanner is Google’s global scale and synchronously replicated relational database. Its research paper was first published in 2012. Spanner…
The System Design Ideas Behind Advanced Search Functions
Search, in its most basic form, can be implemented as a plain inverted index. We build the index as we add/delete/update documents, and use it during search to go from a search term to the documents that contain it. That’s all good. But if you’re more ambitious than building a simplistic lookup by keyword system, read on. This blog post is a gentle introduction to the core ideas behind some of the advanced search functions. By the way, you may recognize that a lot of the ideas here are based on the lucene index.
https://levelup.gitconnected.com/the-system-design-ideas-behind-advanced-search-functions-fa9c7d9010a3
#algorithms #search #design #inverted #index #lucene #scoring #ranking #faceting #mutation
đź’» The byte is not enough
Search, in its most basic form, can be implemented as a plain inverted index. We build the index as we add/delete/update documents, and use it during search to go from a search term to the documents that contain it. That’s all good. But if you’re more ambitious than building a simplistic lookup by keyword system, read on. This blog post is a gentle introduction to the core ideas behind some of the advanced search functions. By the way, you may recognize that a lot of the ideas here are based on the lucene index.
https://levelup.gitconnected.com/the-system-design-ideas-behind-advanced-search-functions-fa9c7d9010a3
#algorithms #search #design #inverted #index #lucene #scoring #ranking #faceting #mutation
đź’» The byte is not enough
Medium
The System Design Ideas Behind Advanced Search Functions
Ever wondered how some of the advanced search functions are implemented? This blog post walks you through the key design ideas behind them.
A Tricky System Design Interview Question: Explain Server Monitoring
A proactive view on how to approach server monitoring during the system design interview rounds. The article is full of unnecessary details, but it still provides a good point on topics like metrics aggregation, time-series databases and pull/push model trade-offs.
https://betterprogramming.pub/a-tricky-system-design-interview-question-explain-server-monitoring-c5be0ce54a30
#systemdesign #infrastructure #sdi #monitoring #metrics #timeseries #aggregation #prometheus #gorilla #scalability #faulttolerance #alerting
đź’» The byte is not enough
A proactive view on how to approach server monitoring during the system design interview rounds. The article is full of unnecessary details, but it still provides a good point on topics like metrics aggregation, time-series databases and pull/push model trade-offs.
https://betterprogramming.pub/a-tricky-system-design-interview-question-explain-server-monitoring-c5be0ce54a30
#systemdesign #infrastructure #sdi #monitoring #metrics #timeseries #aggregation #prometheus #gorilla #scalability #faulttolerance #alerting
đź’» The byte is not enough
Medium
A Tricky System Design Interview Question: Explain Server Monitoring
Test your system design abilities before the next interview
Building an Append-only Log From Scratch
Append-only log, sometimes called write ahead log, is a fundamental building block in all modem databases. It’s used to persist mutation commands for recovery purposes. A command to change the database’s state will first be recorded in such a log before applying to the database. Should the database crash and lose updates, its state can be restored by applying the commands since the most recent snapshot backup. The log’s append-only nature is partly due to functional design — because we only ever need to add new commands to the end — and partly due to performance need — sequential file access is orders of magnitude faster than random file access.
https://eileen-code4fun.medium.com/building-an-append-only-log-from-scratch-e8712b49c924
#design #wal #append #log #distributed #shared #commit #log #database #level #leveldb #compaction #segments #lsmtree
đź’» The byte is not enough
Append-only log, sometimes called write ahead log, is a fundamental building block in all modem databases. It’s used to persist mutation commands for recovery purposes. A command to change the database’s state will first be recorded in such a log before applying to the database. Should the database crash and lose updates, its state can be restored by applying the commands since the most recent snapshot backup. The log’s append-only nature is partly due to functional design — because we only ever need to add new commands to the end — and partly due to performance need — sequential file access is orders of magnitude faster than random file access.
https://eileen-code4fun.medium.com/building-an-append-only-log-from-scratch-e8712b49c924
#design #wal #append #log #distributed #shared #commit #log #database #level #leveldb #compaction #segments #lsmtree
đź’» The byte is not enough
Medium
Building an Append-only Log From Scratch
Real world append-only log features explained by code.
Log Structured Merge Tree (LSM-tree) Implementations (a Demo and LevelDB)
Log structured merge tree, or LSM-tree, is a famous data structure that has been widely adopted by many modern “big data” products, such as BigTable, HBase, LevelDB, etc. Its core idea is very simple and perhaps somewhat counterintuitive if you’re used to the traditional database architecture.
https://eileen-code4fun.medium.com/log-structured-merge-tree-lsm-tree-implementations-a-demo-and-leveldb-d5e028257330
#design #systemdesign #database #lsmtree #leveldb #sstable #index #wal #faulttolerance #performance #scalability
đź’» The byte is not enough
Log structured merge tree, or LSM-tree, is a famous data structure that has been widely adopted by many modern “big data” products, such as BigTable, HBase, LevelDB, etc. Its core idea is very simple and perhaps somewhat counterintuitive if you’re used to the traditional database architecture.
https://eileen-code4fun.medium.com/log-structured-merge-tree-lsm-tree-implementations-a-demo-and-leveldb-d5e028257330
#design #systemdesign #database #lsmtree #leveldb #sstable #index #wal #faulttolerance #performance #scalability
đź’» The byte is not enough
Medium
Log Structured Merge Tree (LSM-tree) Implementations (a Demo and LevelDB)
Background
System Design Interview: Distributed Top K Frequent Elements in Stream
The top k frequent elements question is a classic coding interview question. Viewing through the lens of system design interview however, it suddenly has a special appeal. But before we start our usual system design interview discussion, let’s quickly talk about how it’s solved in the coding interview context, which will facilitate the design discussion that follows.
https://levelup.gitconnected.com/system-design-interview-distributed-top-k-frequent-elements-in-stream-2e92d63d777e
#systemdesign #algorithms #design #topK #frequency #distributed #mapreduce #streaming #kafka #scalability #performance #countminsketch
đź’» The byte is not enough
The top k frequent elements question is a classic coding interview question. Viewing through the lens of system design interview however, it suddenly has a special appeal. But before we start our usual system design interview discussion, let’s quickly talk about how it’s solved in the coding interview context, which will facilitate the design discussion that follows.
https://levelup.gitconnected.com/system-design-interview-distributed-top-k-frequent-elements-in-stream-2e92d63d777e
#systemdesign #algorithms #design #topK #frequency #distributed #mapreduce #streaming #kafka #scalability #performance #countminsketch
đź’» The byte is not enough
Medium
System Design Interview: Distributed Top K Frequent Elements in Stream
A famous coding question posed in system design interview.
System Design Idea: Robust Streaming Data Processing
Streaming data processing, while providing benefits such as freshness and smoother resource consumption compared to their batch counterpart, has historically been associated with disadvantages like being unreliable and having approximate results. Those disadvantages, however, are not inherent characteristics of streaming data processing itself, but rather artifacts of how they have previously been implemented. As shown by Google Cloud Dataflow and the Apache Beam movement in recent years, streaming processing can be just as robust as batch processing. In this blog post, we’ll discuss how to make streaming data processing robust. A lot of the ideas here are based on MillWheel, on which Google Cloud Dataflow is said to be built.
https://levelup.gitconnected.com/system-design-idea-robust-streaming-data-processing-2e9224c33d3f
#systemdesign #infrastructure #design #streaming #beam #dataflow #MillWheel #faulttolerance #reliability #scalability #durability
đź’» The byte is not enough
Streaming data processing, while providing benefits such as freshness and smoother resource consumption compared to their batch counterpart, has historically been associated with disadvantages like being unreliable and having approximate results. Those disadvantages, however, are not inherent characteristics of streaming data processing itself, but rather artifacts of how they have previously been implemented. As shown by Google Cloud Dataflow and the Apache Beam movement in recent years, streaming processing can be just as robust as batch processing. In this blog post, we’ll discuss how to make streaming data processing robust. A lot of the ideas here are based on MillWheel, on which Google Cloud Dataflow is said to be built.
https://levelup.gitconnected.com/system-design-idea-robust-streaming-data-processing-2e9224c33d3f
#systemdesign #infrastructure #design #streaming #beam #dataflow #MillWheel #faulttolerance #reliability #scalability #durability
đź’» The byte is not enough
Medium
System Design Idea: Robust Streaming Data Processing
Ever wondered how streaming processing can be reliable? This article offers a glimpse to the designs adopted by modern cloud services.
đź’ˇ System Design Interview: mini Twitter
How would you design Twitter? It’s essentially the same question as designing Facebook, Linkedin, or many other social media services. The key scenario in question is usually that a user can follow or unfollow other users. Any user can tweet stuff. A user should be able to see tweets, displayed in a certain order, from the users she is following — the so-called timeline. And the core discussion point is designing for scale.
https://eileen-code4fun.medium.com/system-design-interview-mini-twitter-1e0e99bd7377
#systemdesign #sdi #design #infrastructure #twitter #interview #prep #fanout
đź’» The byte is not enough
How would you design Twitter? It’s essentially the same question as designing Facebook, Linkedin, or many other social media services. The key scenario in question is usually that a user can follow or unfollow other users. Any user can tweet stuff. A user should be able to see tweets, displayed in a certain order, from the users she is following — the so-called timeline. And the core discussion point is designing for scale.
https://eileen-code4fun.medium.com/system-design-interview-mini-twitter-1e0e99bd7377
#systemdesign #sdi #design #infrastructure #twitter #interview #prep #fanout
đź’» The byte is not enough
Medium
System Design Interview: mini Twitter
How would you design Twitter? It’s essentially the same question as designing Facebook, Linkedin, or many other social media services. The…