The Distributed Architecture Behind Apache Cassandra
https://blog.gft.com/blog/2017/01/24/the-distributed-architecture-behind-apache-cassandra/
#distributed #system #design #cassandra
💻 The byte is not enough
https://blog.gft.com/blog/2017/01/24/the-distributed-architecture-behind-apache-cassandra/
#distributed #system #design #cassandra
💻 The byte is not enough
GFT Blog English |
The Distributed Architecture Behind Apache Cassandra | GFT Blog English
What's behind the NoSQL database Apache Cassandra? GFT Expert Bruno Tinoco shares some technical information about his journey learning about Cassandra.
The First Few Milliseconds of an HTTPS Connection
http://www.moserware.com/2009/06/first-few-milliseconds-of-https.html
#security #ssl #tls #https #amazon #rsa #handshake
💻 The byte is not enough
http://www.moserware.com/2009/06/first-few-milliseconds-of-https.html
#security #ssl #tls #https #amazon #rsa #handshake
💻 The byte is not enough
Moserware
The First Few Milliseconds of an HTTPS Connection
Convinced from spending hours reading rave reviews, Bob eagerly clicked “Proceed to Checkout” for his gallon of Tuscan Whole Milk and…
Ex-Google engineer shares his experience interviewing people at big tech companies, asking the candidates tough Dynamic Programming questions in particular.
Part 1:
https://medium.com/hackernoon/google-interview-questions-deconstructed-the-knights-dialer-f780d516f029
Part 2:
https://medium.com/@alexgolec/google-interview-questions-deconstructed-the-knights-dialer-impossibly-fast-edition-c288da1685b8
#google #interview #dp #leetcode #problemsolving #algorithms #math
💻 The byte is not enough
Part 1:
https://medium.com/hackernoon/google-interview-questions-deconstructed-the-knights-dialer-f780d516f029
Part 2:
https://medium.com/@alexgolec/google-interview-questions-deconstructed-the-knights-dialer-impossibly-fast-edition-c288da1685b8
#google #interview #dp #leetcode #problemsolving #algorithms #math
💻 The byte is not enough
Medium
Google Interview Questions Deconstructed: The Knight’s Dialer
A Google engineer breaks down a question he used when interviewing people for tech jobs.
Intro To Information Theory
From Bits To The Modern Entropy Function
https://www.setzeus.com/public-blog-post/intro-to-information-theory
#math #cs #informationtheory #shannon #markovchain #entropy #uncertainty #bit
💻 The byte is not enough
From Bits To The Modern Entropy Function
https://www.setzeus.com/public-blog-post/intro-to-information-theory
#math #cs #informationtheory #shannon #markovchain #entropy #uncertainty #bit
💻 The byte is not enough
Setzeus
Intro To Information Theory
The Indonesian Caves of Borneo Island offer an insight to the most primitive form of recorded communication. Around some 40,000 years ago, before the evolution of written language, physical illustrations on cave walls were the most precise method of recorded…
A Deep Dive Into Google BigQuery Architecture
Google’s BigQuery is an enterprise-grade cloud-native data warehouse.
https://panoply.io/data-warehouse-guide/bigquery-architecture/
#google #bigquery #datawarehouse #sql #gcp #dremel #borg #colossus #jupiter #capacitor
💻 The byte is not enough
Google’s BigQuery is an enterprise-grade cloud-native data warehouse.
https://panoply.io/data-warehouse-guide/bigquery-architecture/
#google #bigquery #datawarehouse #sql #gcp #dremel #borg #colossus #jupiter #capacitor
💻 The byte is not enough
Panoply
A Deep Dive Into Google BigQuery Architecture: How It Works [2024 Updated]
Understand how Google BigQuery architecture works. Explore BigQuery best practices for optimizing query performance and providing high cost-effectiveness.
Resumes that made it into FAANG
FAANG style resume tips
https://code.likeagirl.io/resumes-that-made-it-into-faang-f7de9b0c4396
#resume #cv #tips #faang #google #facebook #flipkart #bloomberg #microsoft #servicenow #cisco #jobs
💻 The byte is not enough
FAANG style resume tips
https://code.likeagirl.io/resumes-that-made-it-into-faang-f7de9b0c4396
#resume #cv #tips #faang #google #facebook #flipkart #bloomberg #microsoft #servicenow #cisco #jobs
💻 The byte is not enough
Medium
Resumes that made it into FAANG
There are plenty of blogs available online that teach you how to build a resume. This blog is not about advice regarding building your…
What is Hindley-Milner? (and why is it cool?)
Quick insight on one of the most popular type systems.
https://web.archive.org/web/20181118000004/http://www.codecommit.com/blog/scala/what-is-hindley-milner-and-why-is-it-cool
#hindleymilner #typesystem #scala #haskell #fp #algorithms #typetheory
💻 The byte is not enough
Quick insight on one of the most popular type systems.
https://web.archive.org/web/20181118000004/http://www.codecommit.com/blog/scala/what-is-hindley-milner-and-why-is-it-cool
#hindleymilner #typesystem #scala #haskell #fp #algorithms #typetheory
💻 The byte is not enough
The Netflix Simian Army
Chaos Monkey, a tool that randomly disables production instances to make sure [we] can survive this common type of failure without any customer impact.
https://netflixtechblog.com/the-netflix-simian-army-16e57fbab116
#netflix #chaosmonkey #systemdesign #faulttolerance #cloud #computation #security
💻 The byte is not enough
Chaos Monkey, a tool that randomly disables production instances to make sure [we] can survive this common type of failure without any customer impact.
https://netflixtechblog.com/the-netflix-simian-army-16e57fbab116
#netflix #chaosmonkey #systemdesign #faulttolerance #cloud #computation #security
💻 The byte is not enough
Medium
The Netflix Simian Army
Keeping our cloud safe, secure, and highly available
How many nodes are talked to with Quorum in Cassandra? Also should I use it?
A brief note on Cassandra's consistency levels.
https://foundev.medium.com/cassandra-how-many-nodes-are-talked-to-with-quorum-also-should-i-use-it-98074e75d7d5
#facebook #cassandra #database #systemdesign #consistency #quorum #architecture #design #distributed #keyvalue #dht
💻 The byte is not enough
A brief note on Cassandra's consistency levels.
https://foundev.medium.com/cassandra-how-many-nodes-are-talked-to-with-quorum-also-should-i-use-it-98074e75d7d5
#facebook #cassandra #database #systemdesign #consistency #quorum #architecture #design #distributed #keyvalue #dht
💻 The byte is not enough
Medium
Cassandra: How many nodes are talked to with Quorum? Also should I use it?
This is common early point of confusion with users new to Cassandra, so I just thought I’d drop a brief note in hopes that someone may…
Consistent Hashing
High-level overview of consistent hashing algorithm first published by David Karger et al. back in 1997.
https://tom-e-white.com/2007/11/consistent-hashing.html
#systemdesign #architecture #distributed #system #hashing #consistenthashing #algorithms
💻 The byte is not enough
High-level overview of consistent hashing algorithm first published by David Karger et al. back in 1997.
https://tom-e-white.com/2007/11/consistent-hashing.html
#systemdesign #architecture #distributed #system #hashing #consistenthashing #algorithms
💻 The byte is not enough
Tom White
Consistent Hashing
I’ve bumped into consistent hashing a couple of times lately. The paper that introduced the idea (Consistent Hashing and Random Trees: Distributed Caching Protocols for Relieving Hot Spots on the World Wide Web by David Karger et al) appeared ten years ago…
gfs-sosp2003.pdf
269.5 KB
🧑🔬 Google File System (GFS)
A classical paper on Google File System dating back to 2003 devoted to the distributed file system developed by Google with the main aim to store and process enormous amounts of data at scale effectively.
The work influenced the creation of two other well-known technologies like HDFS and Google's BigTable.
https://www.youtube.com/watch?v=eRgFNW4QFDc
#systemdesign #google #distributed #filesystems #design #reliability #performance #scalability #faulttolerance #research #paper
💻 The byte is not enough
A classical paper on Google File System dating back to 2003 devoted to the distributed file system developed by Google with the main aim to store and process enormous amounts of data at scale effectively.
The work influenced the creation of two other well-known technologies like HDFS and Google's BigTable.
https://www.youtube.com/watch?v=eRgFNW4QFDc
#systemdesign #google #distributed #filesystems #design #reliability #performance #scalability #faulttolerance #research #paper
💻 The byte is not enough
43438.pdf
836.5 KB
🧑🔬 Large-scale cluster management at Google with Borg
An incredible paper on Google's Borg, a Kubernetes predecessor, and how Google successfully managed tens of thousands of machine clusters for over a decade.
Google's Borg system is a cluster manager that runs hundreds of thousands of jobs, from many thousands of different applications, across a number of clusters each with up to tens of thousands of machines.
Watch the Borg presentation at EuroSys 2015:
https://www.youtube.com/watch?v=7MwxA4Fj2l4
#systemdesign #google #borg #distributed #orchestrator #design #reliability #performance #scalability #faulttolerance #research #paper
💻 The byte is not enough
An incredible paper on Google's Borg, a Kubernetes predecessor, and how Google successfully managed tens of thousands of machine clusters for over a decade.
Google's Borg system is a cluster manager that runs hundreds of thousands of jobs, from many thousands of different applications, across a number of clusters each with up to tens of thousands of machines.
Watch the Borg presentation at EuroSys 2015:
https://www.youtube.com/watch?v=7MwxA4Fj2l4
#systemdesign #google #borg #distributed #orchestrator #design #reliability #performance #scalability #faulttolerance #research #paper
💻 The byte is not enough
consistent_hashing_and_random_trees_distributed_caching_protocols.pdf
179.9 KB
🧑🔬 Consistent Hashing and Random Trees: Distributed Caching Protocols for Relieving Hot Spots on the World Wide Web
Another fundamental work on distributed hashing protocols that influenced distributed systems' evolution. Nowadays, consistent hashing is being used in many software distributions like Amazon Dynamo, Apache Cassandra, Riak, Voldemort to name a few. Few big online platforms are known to implement consistent hashing algorithms to scale for performance, availability, and reliability.
More concise take on the matter:
https://www.toptal.com/big-data/consistent-hashing
#systemsdesign #distributed #design #reliability #performance #availability #scalability #research #paper #consistent #hashing
💻 The byte is not enough
Another fundamental work on distributed hashing protocols that influenced distributed systems' evolution. Nowadays, consistent hashing is being used in many software distributions like Amazon Dynamo, Apache Cassandra, Riak, Voldemort to name a few. Few big online platforms are known to implement consistent hashing algorithms to scale for performance, availability, and reliability.
More concise take on the matter:
https://www.toptal.com/big-data/consistent-hashing
#systemsdesign #distributed #design #reliability #performance #availability #scalability #research #paper #consistent #hashing
💻 The byte is not enough
Twine_A_Unified_Cluster_Management_System_for_Shared_Infrastructure.pdf
1.1 MB
🧑🔬 Twine: A Unified Cluster Management System for Shared Infrastructure
If you as me were impressed by the impressive work Google done in their Borg system, and it's successor Kubernetes, then you will definitely enjoy learning how Facebook makes use of their infrastructure in an astonishing paper on Facebook Twine. This tremendous work benefits from experience gained through developing and managing other popular systems like Borg, Kubernetes or Mesos, and aims to scale to more than a million of machines.
https://engineering.fb.com/2019/06/06/data-center-engineering/twine/
#systemdesign #facebook #twine #distributed #orchestrator #design #reliability #performance #scalability #faulttolerance #research #paper #borg #kubernetes #mesos
💻 The byte is not enough
If you as me were impressed by the impressive work Google done in their Borg system, and it's successor Kubernetes, then you will definitely enjoy learning how Facebook makes use of their infrastructure in an astonishing paper on Facebook Twine. This tremendous work benefits from experience gained through developing and managing other popular systems like Borg, Kubernetes or Mesos, and aims to scale to more than a million of machines.
https://engineering.fb.com/2019/06/06/data-center-engineering/twine/
#systemdesign #facebook #twine #distributed #orchestrator #design #reliability #performance #scalability #faulttolerance #research #paper #borg #kubernetes #mesos
💻 The byte is not enough
nsdi13-final170_update.pdf
370.1 KB
🧑🔬 Scaling Memcache at Facebook
While reading Facebook's Twine cluster management system white paper, I noticed an interesting thing in section 5 where paper authors claim to have a highly optimized memcached deployment that can handle "930k lookups per second on an 18-core/36-hyperthread machine". Just think for a second, 930 000 lookup requests per second on a single machine. How the heck they could achieve this kind of performance. To answer this and many other questions, I went on reading another Facebook white paper on optimizing Memcached for meeting the world's largest social network needs.
A short version listing all the big things Facebook incorporated into their version of Memcached can be found here:
https://medium.com/@shagun/scaling-memcache-at-facebook-1ba77d71c082
#systemdesign #facebook #memcached #distributed #caching #design #reliability #performance #scalability #faulttolerance #research #paper #redis #dht
💻 The byte is not enough
While reading Facebook's Twine cluster management system white paper, I noticed an interesting thing in section 5 where paper authors claim to have a highly optimized memcached deployment that can handle "930k lookups per second on an 18-core/36-hyperthread machine". Just think for a second, 930 000 lookup requests per second on a single machine. How the heck they could achieve this kind of performance. To answer this and many other questions, I went on reading another Facebook white paper on optimizing Memcached for meeting the world's largest social network needs.
A short version listing all the big things Facebook incorporated into their version of Memcached can be found here:
https://medium.com/@shagun/scaling-memcache-at-facebook-1ba77d71c082
#systemdesign #facebook #memcached #distributed #caching #design #reliability #performance #scalability #faulttolerance #research #paper #redis #dht
💻 The byte is not enough
atc13-bronson.pdf
729.7 KB
🧑🔬 TAO: Facebook’s Distributed Data Store for the Social Graph
Have you ever wondered how Facebook manages its social graph, containing petabytes of user data? What techniques do they apply to serve billions of reads and millions of writes each second? All this and much more in another great white paper on Facebook's TAO - a geographically distributed data store that provides efficient and timely access to the social graph for Facebook’s demanding workload using a fixed set of queries.
https://engineering.fb.com/2013/06/25/core-data/tao-the-power-of-the-graph/
USENIX ATC '13 - TAO: Facebook’s Distributed Data Store for the Social Graph:
https://www.youtube.com/watch?v=sNIvHttFjdI
#systemdesign #facebook #memcached #tao #socialgraph #api #caching #design #reliability #performance #scalability #faulttolerance #consistency #research #paper #mysql #infrastructure
💻 The byte is not enough
Have you ever wondered how Facebook manages its social graph, containing petabytes of user data? What techniques do they apply to serve billions of reads and millions of writes each second? All this and much more in another great white paper on Facebook's TAO - a geographically distributed data store that provides efficient and timely access to the social graph for Facebook’s demanding workload using a fixed set of queries.
https://engineering.fb.com/2013/06/25/core-data/tao-the-power-of-the-graph/
USENIX ATC '13 - TAO: Facebook’s Distributed Data Store for the Social Graph:
https://www.youtube.com/watch?v=sNIvHttFjdI
#systemdesign #facebook #memcached #tao #socialgraph #api #caching #design #reliability #performance #scalability #faulttolerance #consistency #research #paper #mysql #infrastructure
💻 The byte is not enough
paxos-simple.pdf
92.8 KB
🧑🔬 Paxos Made Simple
In the year 1989 Leslie Lamport, a known computer scientist in the field of distributed systems, published his tremendous work on Paxos — a family of protocols for solving consensus in a network of unreliable or fallible processors. Though from the very beginning, the proposed algorithm was diminished by computer science society due to its complexity, it started gaining significant recognition after almost 10 years since first published having a second coming in 1998. In late 2001, Lamport published a simplified version of the original paper, discarding unnecessary information and providing the description on the backbone of Paxos protocol.
Paxos Simplified by Chris Colohan:
https://www.youtube.com/watch?v=SRsK-ZXTeZ0
#systemdesign #paxos #consensus #lamport #design #reliability #performance #faulttolerance #scalability #consistency #quorum #research #paper #infrastructure #distributed
💻 The byte is not enough
In the year 1989 Leslie Lamport, a known computer scientist in the field of distributed systems, published his tremendous work on Paxos — a family of protocols for solving consensus in a network of unreliable or fallible processors. Though from the very beginning, the proposed algorithm was diminished by computer science society due to its complexity, it started gaining significant recognition after almost 10 years since first published having a second coming in 1998. In late 2001, Lamport published a simplified version of the original paper, discarding unnecessary information and providing the description on the backbone of Paxos protocol.
Paxos Simplified by Chris Colohan:
https://www.youtube.com/watch?v=SRsK-ZXTeZ0
#systemdesign #paxos #consensus #lamport #design #reliability #performance #faulttolerance #scalability #consistency #quorum #research #paper #infrastructure #distributed
💻 The byte is not enough
16cb30b4b92fd4989b8619a61752a2387c6dd474.pdf
186.2 KB
🧑🔬 MapReduce: Simplified Data Processing on Large Clusters
Another seminal work from the past that established distributed systems' evolution for decades ahead. The MapReduce model is probably the most well know programming model designed for processing and generating big data sets with a parallel, distributed algorithm on a cluster.
#systemdesign #mapreduce #design #performance #faulttolerance #scalability #research #paper #infrastructure #distributed #computation #cluster #gfs
💻 The byte is not enough
Another seminal work from the past that established distributed systems' evolution for decades ahead. The MapReduce model is probably the most well know programming model designed for processing and generating big data sets with a parallel, distributed algorithm on a cluster.
#systemdesign #mapreduce #design #performance #faulttolerance #scalability #research #paper #infrastructure #distributed #computation #cluster #gfs
💻 The byte is not enough
Turbine_Facebook’s_Service_Management_Platform_for_Stream_Processing.pdf
1.9 MB
🧑🔬 Turbine: Facebook’s Service Management Platform
for Stream Processing
A scalable service management platform for Facebook’s stream processing service. Turbine is designed to bridge the gap between the capabilities of existing general-purpose cluster management frameworks like Tupperware and Facebook’s stream processing requirements. In production for several years now, Turbine has enabled a boom in stream processing at Facebook.
https://engineering.fb.com/2020/04/21/data-infrastructure/turbine/
#systemdesign #turbine #facebook #streaming #processing #design #performance #acid #reliability #faulttolerance #scalability #research #paper #infrastructure #distributed #cluster #management
💻 The byte is not enough
for Stream Processing
A scalable service management platform for Facebook’s stream processing service. Turbine is designed to bridge the gap between the capabilities of existing general-purpose cluster management frameworks like Tupperware and Facebook’s stream processing requirements. In production for several years now, Turbine has enabled a boom in stream processing at Facebook.
https://engineering.fb.com/2020/04/21/data-infrastructure/turbine/
#systemdesign #turbine #facebook #streaming #processing #design #performance #acid #reliability #faulttolerance #scalability #research #paper #infrastructure #distributed #cluster #management
💻 The byte is not enough
👍1
LogDevice: a distributed data store for logs
A log is the simplest way to record an ordered sequence of immutable records and store them reliably. Build a data intensive distributed service and chances are you will need a log or two somewhere. At Facebook, we build a lot of big distributed services that store and process data. Want to connect two stages of a data processing pipeline without having to worry about flow control or data loss? Have one stage write into a log and the other read from it. Maintaining an index on a large distributed database? Have the indexing service read the update log to apply all the changes in the right order. Got a sequence of work items to be executed in a specific order a week later? Write them into a log, have the consumer lag a week. Dream of distributed transactions? A log with enough capacity to order all your writes makes them possible. Durability concerns? Use a write-ahead log.
https://engineering.fb.com/2017/08/31/core-data/logdevice-a-distributed-data-store-for-logs/
#systemdesign #logdevice #log #wal #logsdb #rocksdb #consensus #paxos #quorum #lsmtree #facebook #storage #design #performance #scalability #research #faulttolerance #infrastructure
💻 The byte is not enough
A log is the simplest way to record an ordered sequence of immutable records and store them reliably. Build a data intensive distributed service and chances are you will need a log or two somewhere. At Facebook, we build a lot of big distributed services that store and process data. Want to connect two stages of a data processing pipeline without having to worry about flow control or data loss? Have one stage write into a log and the other read from it. Maintaining an index on a large distributed database? Have the indexing service read the update log to apply all the changes in the right order. Got a sequence of work items to be executed in a specific order a week later? Write them into a log, have the consumer lag a week. Dream of distributed transactions? A log with enough capacity to order all your writes makes them possible. Durability concerns? Use a write-ahead log.
https://engineering.fb.com/2017/08/31/core-data/logdevice-a-distributed-data-store-for-logs/
#systemdesign #logdevice #log #wal #logsdb #rocksdb #consensus #paxos #quorum #lsmtree #facebook #storage #design #performance #scalability #research #faulttolerance #infrastructure
💻 The byte is not enough
Engineering at Meta
LogDevice: a distributed data store for logs
Visit the post for more.
Database Storage Engines: B-Tree vs LSM-Tree
Have you ever concern yourself with the question of how modern database systems' internal storage works? Well, there are two popular ways to handle data storage — a B-Tree (a generalization of Binary Search Tree) and a Log-Structured Merge Tree, both with having pros and cons.
Here are a few short articles to get a grasp of the trade-offs between the two:
1. https://blog.yugabyte.com/a-busy-developers-guide-to-database-storage-engines-the-basics/
2. https://blog.yugabyte.com/a-busy-developers-guide-to-database-storage-engines-advanced-topics
3. https://rkenmi.com/posts/b-trees-vs-lsm-trees
4. https://tikv.org/deep-dive/key-value-engine/b-tree-vs-lsm/
#systemdesign #dbs #databases #design #btree #lsmtree #storage #performance #sql #nosql #engine
💻 The byte is not enough
Have you ever concern yourself with the question of how modern database systems' internal storage works? Well, there are two popular ways to handle data storage — a B-Tree (a generalization of Binary Search Tree) and a Log-Structured Merge Tree, both with having pros and cons.
Here are a few short articles to get a grasp of the trade-offs between the two:
1. https://blog.yugabyte.com/a-busy-developers-guide-to-database-storage-engines-the-basics/
2. https://blog.yugabyte.com/a-busy-developers-guide-to-database-storage-engines-advanced-topics
3. https://rkenmi.com/posts/b-trees-vs-lsm-trees
4. https://tikv.org/deep-dive/key-value-engine/b-tree-vs-lsm/
#systemdesign #dbs #databases #design #btree #lsmtree #storage #performance #sql #nosql #engine
💻 The byte is not enough