Showing posts with label pcie. Show all posts
Showing posts with label pcie. Show all posts

Wednesday, May 04, 2016

[Links of the day] 04/05/2016 : Openserver Summit & Fortran OpenCoArray


  • OpenCoArray : Fortran is not dead, and the work on the Co array with accelerator demonstrate it.
  • Openserver summit
    • pcie 4.0 : Some really nice improvement with the upcoming standard in term of performance and especially RAS. However not mr-iov capability yet.. This is sorely missing to make PCIe a true contender on the rack scale fabric level. 
    • Azure SmartNIC : Microsoft use FPGA based smartnic to shorten the update cycle of their Azure cloud fabric. Its a really impressive solution. 
    • Persistent Memory over Fabrics : Mellanox pushing for RDMA based persistent memory solution. Probably trying to corner the market quickly as 3dXpoint and Omnipath solution from Intel are just around the corner. However what caught my attention is slide 14:  HGST PCM Remote Access Demo. What is really interesting is that HGST is probably one step away from merging NVM and RDMA fabric onto a single package. With that they would be able to offer a direct competition with DSSD at lower cost ( following the Eth Drive model ). 

Monday, May 02, 2016

[Links of the day] 02/05/2016 : All about storage @ Intel IDF 16 + no more secret


Tuesday, March 01, 2016

[Links of the day] 01/03/2016 : DSSD , Datacenter design [book] and latent faults [paper]

  • DSSD : EMC released into the wild DSSD product (acquired last year). Quite a beast: all flash, 10 M IOPs, 100μs latency, 100GB/s BW, 144TB/5U. It use a PCIe fabric to connect the storage to the compute nodes, however I expect them to move soon to infiniband / omnipath fabric based on the talk they recently made.
  • Datacenter Design and Management: book that surveys datacenter research from a computer architect's perspective, addressing challenges in applications, design, management, server simulation, and system simulation.
  • Unsupervised Latent Faults Detection in Data Centers : talk and paper that look at automatically enable early detection and handling of performance problems, or latent faults. These faults "fly under the radar" of existing detection systems because they are not acute enough, or were not anticipated by maintenance engineers.
Rolex Deep Sea Sea Dweller (DSSD)


Wednesday, November 18, 2015

Links of the day 18/11/2015: DHT comparison, decentralized control plane, NVMe for mobile

  • Xenon : set of software components and a service oriented design pattern delivering a decentralized Control Plane Framework by Vmware
  • DHT comparison : paper analyzing churn , scale , performance of the 4 main DHT under churn.
  • PCIe/NVMe for mobile : while the argument might make sens the slide deck carefully avoid the elephant in the room : what is the power / performance cost ?


Wednesday, September 16, 2015

The upcoming Storage API battle

There is an interesting trend within the storage ecosystem. We are witnessing a polarization of the offer. On one side, we are seeing the rise of high performance rack scale solution(DSSD, NVMe over fabric solution, etc..) . And on the other side we have the object storage solution which are more datacenter scale. While both leverage heavily non volatile memory they play a different different role within the ecosystem.

The rack scale storage target “very” high performance solution, delivering very low latency high bandwidth access time. Often in the 100 of usec or less. However often these solution come at a higher financial cost due to more expensive hardware (custom NVM), network fabric (IB+NVMe, PCIe + NVMe, Omnipath +NVMe, Pure PCIe, etc..) and significant power consumption (>2000W/5U for DSSD). Finally, these offer access via a specialized API that needs to be either accessed natively or adapted to other more standard one.
On the other side we have the object storage solution. Users access object storage through applications that will typically use a REST API. This makes object storage ideal for all online, Cloud environments. Moreover they tend to be a lot more cost efficient especially with the rise of Ethernet connected drives (up to 50% less TCO).
Stuck in the middle is the classic Filer / POSIX compliant solution that seems to slowly dwindle away. To a certain extent the rack scale solution should have a bright future in the niche (but still significant) market for enterprise that still consider that their application requires custom for what they think is a custom problem. On the other side the object storage is gaining momentum by riding the unstoppable cloud tide.



While, both technology can and should co-evolve, they both suffer from software limitation and to a certain extend hardware one as bandwidth and latency get dangerously close to what cpu are capable to handle. This requires a drastic shift on how applications are developed if user want to actually get any benefit from these solution. However, few company are willing to risk specializing their code using an API that can become obsolete when the next generation of storage solution pops up.
Storage startup/company out there start to discover that  API is playing a significant role success of their product and performance while still important will lose of its importance. It will either make rewriting applications to access your storage infinitely easier task or transform it into a painful experience by forcing to go through hoops and/or adaptation layers with the performance cost associated. 
The fight for the next generation storage API only started. There is and will be more push toward standardization which will be fueled by the customer tiredness with every revolving siloed point solution. People who use object storage one want it to behave more like POSIX storage, but they also want to keep the storage costs at an object level and improve the performance. On the other hand people using rack scale storage want to retain the performance but increase its simplicity and also want the price to come down. It is going to be extremely hard to deliver both but hopefully we might finally see a rationalization of the storage market as having an object storage system that allows byte range access is very appealing. 

Thursday, August 27, 2015

Links of the day 27/08/2015 : PCIe fabric, One Time SSH key and China ARM processor

  • Dolphin PCIe fabric : seems that the PCIe fabric is getting a lot of news recently, nice to see that we are reaching the 300 nanoseconds latency. While greater bandwidth and lower latency is really great this bring a certain amount of challenge as DRAM access time is still around 10-30 nanoseconds. Implying that if we the CPU needs to do anything with this data it only has ~250 nanoseconds to do so before moving to the next one. I wonder when we wills start to see event based GPGPU or FPGA approach in CPU in order to handle such performance. Maybe its time to move toward stream oriented CPU architecture ? 
  • One-Time SSH Keys : I like this concept, it allow the use of single-use key ready for connecting to untrusted system. With single-use keys, even if the key is compromised, it will have already been used, and would be of little use.
  • China ARM : Homegrown ARM v8 based cpu chips. Seems that the Chinese are getting serious in leveraging homegrown solution in the CPU field and especially HPC. From this expertise they hope to spin off in the scale out computing (high density, low power). Maybe the ARM server breakthrough will come from the Asian market? 


Wednesday, August 26, 2015

Links of the day 26/08/2015 : 3Dx point & SAP HANA , PCIe fabric, Hashmap for NVM

  • PCIe switch : after buying PLX avago is trying to ramp up the PCIe fabric business. The interesting bit will come up when it will start to compete against Intel omni path fabric. While PCIe fabric provide the advantage of direct device access and MR-IOV is really attractive a less HW and more Network oriented fabric might have certain advantage. Maybe we will start to see true heterogeneous Clos topology fabric in the future. PS: bonus patent for a PCIe fabric using non transparent bridge.
  • 3Dx point and SAP HANA : while the post in itself is not that interesting ( enable bigger better in memory system). The responses from Hasso Plattner to it show a rather puzzling perspective, it seems that there is a disconnection between the "marketing" and actual real use case. HANA can "theoretically" be deployed on massive machine but in reality you only need a third of that. Basically very , I repeat, VERY few (to no) companies need that type of specs. Maybe, if SAP started to offer modular HANA solution over their clouds, customers will finally avoid the massive headache that is sizing (and adopt the product more easily). By modular I mean at providing Memory re-sizing at runtime with pay for what memory you consume model. 
  • NVC-Hashmap : A Persistent and Concurrent Hashmap For Non-Volatile Memories

Thursday, May 14, 2015

Links of the day 14 - 05 - 2015

Today's links 14/05/2015 : #Openstack #container , Http/2 PCie switch, #google Heracles resources management
  • magnum : Openstack container project in all its glory .. more tools and knobs ( but sadly no production ready distribution without heavy work)
  • HTTP/2 : how architecting Websites will change with the coming HTTP/2 Era 
  • Avago PEX9700 : the first PCIe switch product release after the acuqisition of PLX by avago. Some nice feature, host to host coms, Nic mode dma, etc.. 
  • Heracles : Descendant of Google pegasus project, moving from power management to full blown resources orchestration.





Tuesday, May 05, 2015

The rise of micro storage services


Current and emergent storage solutions are composed of sophisticated building blocks: dedicated fabric, RAID controller, layered cache, object storage etc. There is the feeling that storage is going against the current evolution of the overall industry where complex services are composed of small, independent processes and services, each organized around individual capabilities.

What surprised me is that most of the storage innovation trends focus on very high level solutions that try to encompass as many features possible in a single package. Presently, insufficient efforts are being made to build storage systems based on small, independent, low-care and indeed low-cost components. In short, nimble, independent modules that can be rearranged to deliver optimal solution based on the needs of the customer, without the requirement to roll out a new storage architecture every time is simply lacking - a "jack of all trade" without,or limited "master of none" drawbacks or put another way modules that extend or mimic what is happening in the container - microservices space.

Ethernet Connected Drives

Despite this, all could change rapidly as an enabler (or precursor, depending how you look at it), of this alternative solution as it is currently emerging and surprisingly, coming from the Hard Drive vendors : Ethernet Connected Drives [slides][Q&A].This type of storage technology is going to enable the next generation of hyperscale cloud storage solution. Therefore, massive scale out potential with better simplicity and maintainability,not to mention lower TCO.

Ethernet Connected Drives are a step in the right direction as they allow a reduction in capital and operating costs by reducing:
  • software stack (File System, Volume Manager, RAID system);
  •  corresponding server infrastructure; connectivity costs and complexity; 
  • granularity which enable greater variable costs by application (e.g. cold storage, archiving, etc.).
Currently, there are two vendors offering this solution : Seagate with Kinetic and HGST with the Open Ethernet Drive. In fact we are already seeing some rather interesting applications of the technology. Seagate released a port of SheepDog project onto its Kinect product [Kinect-sheepdog] there by enabling the delivery of a distributed object storage system for volume and container services that doesn't requires dedication. Indeed there is a proof of concept presented HEPiX of HGST drive running CEPH or Dcache. While these solutions don’t fit all the scenarios, nevertheless, both of these solutions demonstrate the versatility of the technology and its scalability potential (not to mention the cost savings).

What these technologies enables is basically the transformation of the appliances that house masses of HDD into switches thereby eliminating the need for a block or file header as there is now a straight IP connectivity to the drive making these ideal for object based backends.

Emergence of fabric connected Hardware storage:

What we should see over the next couple of years is the emergence of a new form of storage appliance acting as a fabric facilitator for a large amount of compute and network enable storage devices. To a certain extend it would be similar to HP's moonshot except with a far greater density.

Rather than just focusing on Ethernet, it would be easy to see PCI, Intel photonic, Infiniband or more exotic fabrics been used. Obviously Ethernet still remains the preferred solution due to its ubiquity in the datacenter. However, we should not underestimate the need for a rack scale approach which would deliver greater benefit if designed correctly.
While HGST Open Ethernet solution is one good step towards the nimble storage device, the drive enclosure form factor is still quite big and I wouldn't be surprised if we see a couple of start-ups coming out of stealth mode in the next couple of months with fabric (PCIe most likely) connected Flash. This would be an equivalent of the Ethernet connected drive interconnected using a switch + backplane fabric as shown in the crudely designed diagram below.







Is it all about hardware?

No, indeed quite the opposite. That said, there is a greater chance of penetration of new hardware in the storage ecosystem as compared to the server market. This is probably where ARM has a better chance of establishing a beach head within the hyperscale datacenter as the microserver path seems to have failed.
What this implies is that it is often easier to deliver and sell a new hardware or appliance solution in the storage ecosystem than a pure software one. Software solutions tend to take a lot longer to get accepted, but when they pierce through, they quickly take over and replace the hardware solution. Look at the object storage solution such as CEPH or other hyper-converged solution. They are a major threat to the likes of Netapp and EMC.
To get back on the software side as a solution, I would predict that history repeats itself to varying degrees of success or failure. Indeed, like the microserver story we see, hardware micro storage solutions while rising, at the same time we see the emergence of software solutions that will deliver more nimble storage features than before.

In conclusion, I feel that we are going to see the emergence of many options for a massive scale-out, using different variants of the same concept: take the complex storage system and break it down to its bare essential components; expose each single element as its own storage service; and then build the overall offer dynamically from the ground up. Rather than leveraging complexed pooled storage services we would have dynamically deployed storage applications for specific demands composed of a suite of small services, each running in its own process and communicating with lightweight mechanisms.These services are built around business capabilities and independently deployable by fully automated deployment machinery. There is a minimum of centralized management relating to these services, which may be written in different programming languages and use different data storage technologies, . which is just the opposite of current offers where there are a lot of monolithic storage applications (or appliances) that are then scaled by replicating across servers.

This type of architecture would enable a true, on-demand dynamic tiered storage solution. To reuse a current buzzword, this would be a “lambda storage architecture”.
But this is better left for another day’s post that would look into such architecture and lifecycle management entities associated with it.





Tuesday, April 28, 2015

Links of the day 28 - 04 - 2015

Today's links 28/04/2015: #Rump Kernel Stack, Disque Distributed In Memory MQ, #FusionIO new PCIe product, Power level estimation of VM systems
  • Ramp Stack : Nginx, MySQL, and PHP built on Rump Kernels without rearchitecting the application. Most of the work requires the app to be cross compile correctly (Nginx & MySQL). This implies that Unikernel-compatible unmodified POSIX C and C++ applications “just work” on top of Rump Kernels, provided that they can be cross- compiled.
  • Disque :a distributed, in memory, message broker by Redis folk. Not production ready but a promising start.
  • Pcie Flash : fusion IO is still kicking and deliver an interesting solution: up to 350,000 I/O operations per second (IOPS) on random reads and 385,000 IOPS on random writes (on the 3.2 TB model) with a 15k nanosecond write latency and 2.8 GB/sec of read bandwidth. .. However I still don't get why they don't want to use NVMe tech
  • Process-level Power Estimation in VM-based Systems : the authors describe a fine-grained monitoring middleware providing real-time and accurate power estimation of software processes running at any level of virtualization in a system. 




Thursday, April 23, 2015

Links of the day 24 - 04 - 2015

Today's links 24/04/2015: iOS fail, Memory Snooping Protection in HW, NVMe/PCIe fabric for storage, High performance server application framework
  • Back to the Future With C++ and Seastar : Meetup notes on Seastar, open source server application framework written in C++ that presents a future/promise based API with 7 million requests per second served on a single machine. [slides]
  • No iOS zone:  PoC of attack : within WiFi hotspot range iOS devices are rendered unstable / unusable by constant reboots.
  • Cloud security reaches silicon : Hardware implementation of method for thwarting memory snooping or inference across VM attacks by disguising memory-access patterns
  • NVMe/PCIe as a Storage Interface : an very good overview of the market, future solution and form factors.

Friday, February 20, 2015

Links of the day 20 - 02 - 2015

Links of the day 20/02/2015:Probabilistic counters, Rack PCIe Fabric, Complex Networks robustness, High performance Network protocol
  • Characterizing Storage Workloads with Counter Stacks : Really great approach to use hyperloglog algorithm to characterize workloads. It allow the identification of working set sizes without the large memory overhead . The key advantage is it allow the generation of accurate MRC for large workload without the tremendous overhead which then can be used to tweak the cache size and for online placement decision.
  • Rack Level PCIe fabric : looks like the competition is heating up for the next gen rack level fabric. There is some really nice feature in there, advanced topology ( torus, 3D fat tree), support for native tcp/ip , rdma and native. And last but not least sharing of I/O by the assignment of the VFs of SR-IOV . Not to forget that most ops are sub micro seconds. Its nice to see the rack level fabric competition heating up.
  • Improving the Robustness of Complex Networks with Preserving Community Structure : 3-step strategy to improve the robustness of a network, while retaining its community structure, and also its degree distribution.
  • Trickles : Stateless High Performance Networking that relies on Transport continuations information within the packet for congestion control algorithm , effectively removing the state maintenance on both side. Creating a stateless protocol at the cost of periodical update signal.

Wednesday, February 11, 2015

Links of the day 11 - 02 - 2015

Today's links 11/02/2015: Test and optimization articles, scaling product team, PCIe vs Eth , Distributed Sys fallacies
  • 100 Must-Read Articles on Testing and Optimization : data driven, big data, a/b testing etc.. The best articles from 2014
  • Scaling a product team : lesson learned from Intercom on how they scaled a product building team, and the nitty gritty involved in getting valuable product out the door as fast as possible.
  • Eight Fallacies of Distributed Computing : very good tech talk with real life encounter of the fallacies.
  • PCIe vs Ethernet : with the rise of Intel’s silicon photonics (SiPh) optical PCIe (OPCIe) and other PCIe fabric, is it time to fragment your datacenter and use fast PCIe rack fabric and Eth for cross rack one. To be honest time will tell as you already know the best technology doesn't always win.