Showing posts with label reliability. Show all posts
Showing posts with label reliability. Show all posts

Thursday, November 08, 2018

[Links of the Day] 08/11/2018 : large scale study of datacenter network reliability, What to measure in production, Failure Mode effect Analysis

  • A Large Scale Study of Data Center Network Reliability : the authors study reliability within and between Facebook datacenters. One of the key findings is the growth in complexity, heterogeneity and interconnectedness of datacenter increase the rate of occurrence of unwanted behaviours. Moreover, this seems to be also a key potential limiting factor for world scale spanning infrastructure undergoing rapid organic growth.
  • Understanding Production: What can you measure? : what do you need to monitor and measure in production. Very good summary of many blog post out there.
  • Failure Mode Effects Analysis (FMEA) : once you start reaching a certain production scale and more stringent requirement kicks in ( unless you were unlucky enough to have them at the get-go). You might want to run a failure modes and effects analysis (FMEA) is a step-by-step approach for identifying all possible failures in a design, a manufacturing or assembly process, or a product or service. While it was mainly designed to address shortcomings in the manufacturing industry, it is still extremely useful for IT system analysis, especially when you want to prepare yourself pre-rollout of a chaos monkey like system.


Wednesday, October 19, 2016

[Links of the day] 19/10/2016 : #AI hard problems, Dark Silicon & Reliability , Transport Layer Dev Kit

  • Applied AI hard problems : current and future AI hard problem, the interesting bit is the "emergent" behavior aspect that computer scientist are trying to achieve. Where AI is not tailored for a specific problem by adapt to the environment it encounter. 
  • Dark silicon & Hardware Reliability : the authors look at the impact of the dark silicon approach ( when not all component are turned on when the system is up) and how to leverage the "dark" ratio to maximise lifespan of hardware. [slides]
  • TLDK : project lead by Intel within the fd.io framework. It is trying to adresse the lack of high level ( as in layer 4 ) packet processing capabilities. The project aim at delivering UDP/TCP etc.. packet processing on top of vector packet processing of FD.io (which can works on top of DPDK). By doing so Intel will be able to finally have a comprehensive framework which will enable DPDK based solution to flourish beyond the pure networking stack (NFV) solution.

Wednesday, April 13, 2016

[Links of the day] 13/04/2016: distributed system debugging tool, SRE book note, resilient ad serving at scale

  • Shiviz : tools for debugging distributed system at scale via helping visualization of logs generated throughout the cluster [repo]
  • Notes on Google's Site Reliability Engineering book : this provide a good overview of the content of the book and provide enough details that can be used to research each aspect / chapter independently.
  • Resilient ad serving at Twitter-scale : its interesting to see that their is a correlation between query latency and revenue  for ad serving. This stem from the fact that latency for answering an ad query is dependent of the number of participant in the auction,and obviously the more participant the higher the revenue. However with the increased latency the higher the risk is to time out and hence revenue loss. Twitter use an adaptative system in order to maximise revenue while maintaining resiliency (availability), scalability, resource-utilization.