- A Large Scale Study of Data Center Network Reliability : the authors study reliability within and between Facebook datacenters. One of the key findings is the growth in complexity, heterogeneity and interconnectedness of datacenter increase the rate of occurrence of unwanted behaviours. Moreover, this seems to be also a key potential limiting factor for world scale spanning infrastructure undergoing rapid organic growth.
- Understanding Production: What can you measure? : what do you need to monitor and measure in production. Very good summary of many blog post out there.
- Failure Mode Effects Analysis (FMEA) : once you start reaching a certain production scale and more stringent requirement kicks in ( unless you were unlucky enough to have them at the get-go). You might want to run a failure modes and effects analysis (FMEA) is a step-by-step approach for identifying all possible failures in a design, a manufacturing or assembly process, or a product or service. While it was mainly designed to address shortcomings in the manufacturing industry, it is still extremely useful for IT system analysis, especially when you want to prepare yourself pre-rollout of a chaos monkey like system.
A blog about life, Engineering, Business, Research, and everything else (especially everything else)
Showing posts with label chaos. Show all posts
Showing posts with label chaos. Show all posts
Thursday, November 08, 2018
[Links of the Day] 08/11/2018 : large scale study of datacenter network reliability, What to measure in production, Failure Mode effect Analysis
Labels:
analysis
,
chaos
,
datacenter
,
links of the day
,
measure
,
monitoring
,
network
,
production
,
reliability
Tuesday, September 29, 2015
Links of the day 29/09/2015 : Time series in SQL, Chaos engineering, Distributed Systems Architecture
- Storing Time series in SQL : how to efficiently store time series data in postgres SQL or other SQL db
- Principles of Chaos Engineering : discipline of experimenting on a distributed system in order to build confidence in the system’s capability to withstand turbulent conditions in production.
- Architectural patterns of resilient distributed systems : while this is a good talk about resiliency I have a slight issue with the term itself. I have a preference for ductile rather than resilient. As often system that are designed to be resilient are extremely brittle when the reaching the breaking point [video] [slide-deck]
Labels:
architecture
,
chaos
,
Distributed systems
,
links of the day
,
sql
,
time series
Subscribe to:
Posts
(
Atom
)

