Showing posts with label arm. Show all posts
Showing posts with label arm. Show all posts

Sunday, November 22, 2020

HPC ecosystem - SC20

 This article is a quick summary of SC20 trends and the current state of the HPC ecosystem from a tech and market perspective.


Technology-wise there are three main competing HPC architectures: 

  • Commodity (e.g. Intel)
  • Commodity + accelerator (e.g. GPUs)
  • Lightweight cores (e.g. IBM BG, Xeon Phi, TaihuLight, ARM )


Commodity systems represent the bulk of the systems out there. However, commodity + accelerator are ramping up their presence aggressively. Nvidia dominates this market segment with 142 systems out of 149. With Intel scooping 4 with it's Phi solution. Lightweight cores systems are a minority with only four systems. But with the new A64FX and a renewed appetite for custom chips, this might change rapidly.

Intel is still dominating the ecosystem, with 92% of the shares, followed by AMD with 4%. However, this might change rapidly with AMD EHP technology ramping up. Another aspect is that AMD technology tends to be more open-source friendly, which can make it more attractive long term. Not to mention that their GPU also start to become highly competitive in the AI space.





From a market size, the HPC market was $39.0 billion in 2019, up 8.2% from $36.1 billion in 2018. Predictions show growth to $55.0 billion in 2024. Most of the growth was led by government spending after six years of growth led by industry. The number of system in industry vs public is not equally divided with ~50% each.





One notable change is double-digit growth of cloud HPC related market. Cloud grew 17.8% to $1.4 billion; however, this might only be the tip of the iceberg as many companies might be using HPC like system in the cloud without labelling it as HPC. Cloud solutions are heavily displacing low-end HPC segment. Entry and mid-range level server classes have the slowest growth in years as consumers prefer to buy HPC as a service solution and reduce their CAPEX. 







AI is still heavily influencing the HPC infrastructure market as it represents a considerable opportunity for HPC solution vendors. HyperscaleAI infrastructure by itself is about $8 billion. It seems that for the moment, AI and HPC future are closely intertwined.


Sources: Intersect360 research - Pre-SC20 Market Update & Jack Dongarra - An overview of HPC

Sunday, November 01, 2020

ARM ecosystem disintegration and the rise of RISC-V

#ARM acquisition by #Nvidia is making people uneasy. 

And the early sign of the unravelling of the #ARM ecosystem start to appear: ThunderX3 general-purpose ARM CPU has been cancelled.

One would ask why spending $$ to build a better product and increase its number of consumers if, for that, it will have to use the Nvidia IP and compete directly against the IP owner.
If you combine this with the difficult viability of putting together a general-purpose #ARM alternative to #Intel / #AMD as #ARM vendors are effectively competing on cost with much lower volumes.

We start to understand why Marvell decided to shift toward the much more trendy IPU/PDU/Smartnic market.

On the other hand, I think we will see an acceleration of RISC-V adoption. Eating away at the traditional #ARM market share. This will be driven by the large scale edge deployment of #riscv sees chips with a RISC-V core and an #NPU (neural processing unit). These chips can be churned out at incredibly cheap cost, less than $10, and these will become ubiquitous really rapidly.

It might take 10-15 years but ultimately this will seal the fate of the ARM franchise.




Monday, June 22, 2020

Stop throwing GPU at HPC, Not all scientific problems are compute dense

The current race to exascale has put a heavy emphasis on GPU-based acceleration at the detriment of other HPC architecture. However, Crossroads and Fugaku supercomputer are demonstrating that it is not all about GPU.



The vast majority of the (pre-)exascale machines are relying heavily on GPU acceleration targeting scientific problems that can be cast as dense matrix-matrix multiplication problems.

However, there are large numbers of scientific problems that are not compute dense. And such GPU architectures are ill-equipped to accelerate these problems. Sadly, the current trends seem to have relegated those type of scientific challenges to second class citizens in the HPC world. If you look at extreme-scale graph problems by example, the graph500 benchmark clearly shows that these type of problem have been orphaned. 4 out of ten systems are more than seven years old and nearing their end of life. Moreover, the newer systems show marginal progress toward accelerating extreme-scale graph traversal. 

I understand that the current machine learning hype heavily influences the HPC ecosystem. However, we have to remind ourselves that there is life beyond FLOPS. And the Fugaku and Crossroads system demonstrates it is possible to achieve strategic compute leadership without sacrificing the architecture to the altar of exaflop compute dense gods. 



The Japanese latest ARM-based Fugaku supercomputer is demonstrating that it can address both compute dense GPU optimised and the one that not reducible to dense linear algebra and therefore incompatible with GPU technologies. The Japanese supercomputer built around the ARM v8.2A A64FX CPU just picked up the number one in the HPC Top 500 Green benchmark and the Graph500 BFS benchmark.

Hopefully, this will be a wake-up call within the HPC community to properly fund R&D efforts orthogonal to the compute dense and exaflops benchmark friendly architecture.


Update 22/06/2020 : after publishing this article Fugaku just got ranked Nb 1 at the top 500 Linpack benchmarks with close to half an exaflops! (415 Petaflops).And Fugaku is pretty much topping every single HPC ranking :

  • Top500 : Nb 1.
  • Top500-Green : Nb 1.
  • HPCG : Nb 1.
  • HPL-AI: Nb 1.
  • Graph500 : Nb 1.







Friday, October 06, 2017

[Links of the Day] 06/10/2017 : HPC routing topology on dependency graph, Arm network stack development, Multicore graph processing

  • Routing on the Channel Dependency Graph : the author aim are providing a toolset for topology calculation for HPC network. 
  • Arm for network stack developers : arm is trying to slowly move up the stack into the data centre world. For that, it needs to address one of its main limitation: IO. This slide deck describe the current effort to tackle network stack limitation ( and support RDMA ) as well as providing pointer where ARM or devs can push the Linux stack further. As Linux is pretty much the only viable software stack for such hardware infrastructure in the datacenter. 
  • Multicore graph processing : this present a very good overview of the Multicore graph processing problem, solution landscape and where it's heading. Graph problems are essential to solve in the domain of social network modelling as well as for items recommendation or search and website ranking. [slides]



Friday, October 14, 2016

[Links of the day] 14/10/2016 : Docker Infrakit , erasure Code for big-data and ARM research summit

  • InfraKit : docker answer to public cloud lock in. It allow devs to easily deploy their systems on various cloud infrastructure without code change. 
  • Erasure Coding for Big-data Systems : Phd Thesis of Rashmi Korlakai Vinayak on erasure code for very large data systems. The author analyse the requirement and provide potential solution allowing resource efficient distributaire storage codes . The authors looks at the various trade-off that can be used to guarantee durability while limiting ressource usage. 
  • ARM Research Summit 2016 : live blog of the keynotes, a lot of the research issue are similar to x86 one. Which can be worry some as ARM needs to be able to differentiate itself from Intel especially in the server market.

Wednesday, June 01, 2016

[Links of the day] 01/06/2016 : Megaprocessor, ARM A73 Artemis and AsAP project

  • Megaprocessor : Somebody decided to build a processor using actual individual transistor component. This impressive both in scale and dedication. [videos ]
  • Artemis : The ARM Cortex A73 has been released. It is the 64-bit successor to the Cortex A17 . Yes,this is confusing, the naming and release cycle doesn't really indicate the actual generation / functionality. It is clearly aimed at the smartphone market with its little/big core architecture and is a marvelous piece of tech. 
  • AsAP : Asynchronous Array of Simple Processors , remind me of the connection machine architecture. Except that it use a single-chip processing system comprised of a large number of fine-grain asynchronously-operating programmable processors connected by a reconfigurable network.

Saturday, March 26, 2016

PureStorage bring us one step closer to micro storage architecture

Pure storage just released its Flashblade product. It is an fabric connected object storage solution. It is a modular solution composed of a large numbers of blades which are each made of :
  • 8TB to 52TB raw NAND storage capacity : a lot but still take less than half the real estate space on each blade.
  • NV-RAM+supercapacitor write buffer : when your NAND is still too slow you want to have a persistent buffer of NVRAM to handle the bursts
  • ARM CPU + FPGA : to deal with the “low level” operations such as erasure code, etc..
  • 8 core Xeon System on chip : for moving the computation to where the data is located, pretty much all the high level operation such as NFS , S3 , object storage etc.. 
  • 40 Gbit ethernet : that s where the data gets out
  • PCIe fabric networking : in chassis solution linking compute, storage cards via a proprietary protocol, what’s interesting is that the system is self contained and scaling with other box goes through the 10 Gb/s connectivity and not a proprietary fabric link. Which implies that it doesn’t need exotic solution once you go past the box boundaries. This is great as it makes it easy (and cheap) to scale however I wonder what are the implication in term of performance once you start crossing boundaries.


What is interesting is that, when you look at Purestorage solution, they decided to integrate the high level compute aspect of storage directly with the low level one in a single blade. They ended of with an hybrid solution combining ARM and FPGA for low level aspect such as deduplication, erasure code. And the Xeon for the object storage and file system solution. 

One can assume that the decision behind such architecture was driven by the customers requirement that tend to want a high performance Jack of all trade solution. I can picture the product manager arguing for supporting every scale out storage protocol popular at the moment. However, Jack always end up master of none and to over compensate PureStorage had to pump up the compute capabilities.

While this seems like a good choice it is also counter productive in term of Watt per GB coupled with a lot of real estate wasted or duplicated. Don’t get me wrong, what Pure achieved with the flashblade is impressive but I can’t stop thinking that they should have taken it a step further.

This type of high performance, high-cost and high-power architecture technology is a right step toward micro storage architecture which delivers low cost low power high performance and scalability features. Now it is all about trimming down the system while maintaining scalability by dividing the blade system into a much larger number of smaller nodes, literally offering what the ethernet connected equivalent of HGST with flash.

However this might also implies that you won’t be able to offer support for every single storage solution out there (NFS, S3, block, etc..) without having to rely on either client side processing or using a frontend. This should be achievable while maintaining excellent performance, the key to this will hide in the detail of the core storage api employed.

Monday, January 18, 2016

[Links of the day] 18/01/2016 : Post quantum crypto, ARM virtualization, Scalable C

  • Post-quantum key agreement : First, 99% of people out their do not use encryption correctly, so you should not be worried about post quantum crypto because you are not protecting yourself in the pre-quantum era. Now if you are part of the 1%, bad news you key size just jump from a couple of Kbytes to Mbytes (and maybe GB)... Very good read explaining the challenges ahead and existing gap of crypto solution in the upcoming post quantum era.
  • ARM virtualization extensions : In depth look at the ARM virtualization feature. Maybe we can see a glimpse of what can be done with the AMD ARM server push. 
  • Scalable C : book on how to make C scalable by the founder of ZeroMQ. Some good things, some bad one, a lot of grief toward C++. I personally love the C language but sometimes pitting one language against another doesn't help without context. Pick the tools that suits best and sometimes yes it means picking the one that makes collaboration efficient rather than make the code efficient. 

Thursday, January 07, 2016

Links of the day 07/01/2016: DoD meet cloud, ACM queue on NVM , ARM v8 evolution

  • ARM v8 evolution : what is happening and how 
  • Non-volatile Storage : impact of NVM on the datacenter, and ACM queue article. If you follow the links of the day there is not much surprise in this article but its a good summary of what is happening out there.
  • When DoD meet cloud : document describing the impact and needs created by the adoption of cloud services on governmental organisations : DoD Needs an Effective Process to Identify Cloud Computing Service Contracts

Thursday, November 19, 2015

Links of the day 19/11/2015: #ARM rack scale infrastructure fabric, Canonical ByoHW #Cloud software and Next Gen #Bigdata Storage

  • X-Tend : Applied micro get into the rack scale computing market with its ARM server and now the fabric to enable and support dis-aggregation of the resources. What is interesting there is since the ARM processor are more nimble in a certain sens it might make them more suitable for this paradigm than the big Intel processor. And if you throw into the mix hybrid and heterogeneous processors might end up having a better fit as you can more easily tailor the resource to your need as well as the power consumption coupled with it.
  • Autopilot : canonical private ByoHW cloud solution (more about that in another post).
  • Pyro: A Spatial-Temporal Big-Data Storage System for high resolution geometry queries and dynamic hotspots.


Friday, October 02, 2015

Thursday, August 27, 2015

Links of the day 27/08/2015 : PCIe fabric, One Time SSH key and China ARM processor

  • Dolphin PCIe fabric : seems that the PCIe fabric is getting a lot of news recently, nice to see that we are reaching the 300 nanoseconds latency. While greater bandwidth and lower latency is really great this bring a certain amount of challenge as DRAM access time is still around 10-30 nanoseconds. Implying that if we the CPU needs to do anything with this data it only has ~250 nanoseconds to do so before moving to the next one. I wonder when we wills start to see event based GPGPU or FPGA approach in CPU in order to handle such performance. Maybe its time to move toward stream oriented CPU architecture ? 
  • One-Time SSH Keys : I like this concept, it allow the use of single-use key ready for connecting to untrusted system. With single-use keys, even if the key is compromised, it will have already been used, and would be of little use.
  • China ARM : Homegrown ARM v8 based cpu chips. Seems that the Chinese are getting serious in leveraging homegrown solution in the CPU field and especially HPC. From this expertise they hope to spin off in the scale out computing (high density, low power). Maybe the ARM server breakthrough will come from the Asian market? 


Tuesday, March 17, 2015

Links of the day 17 - 03 - 2015

Today's links 17/03/2015: hash table datastructure and hashing Algorithm, #ARM virtualization, Hyper-converged Solution

from Adexchanger comic strip

Thursday, February 05, 2015

Links of the day 05 - 02 - 2015

Today's links 05/02/2015 : #openstack #neutron with #dpdk , #ARM A72 coherency, Captain proto
  • Openstack* Neutron Accelerated by DPDK : dpdk is gaining a lot of momentum, to bad that they decided to adopt their own API model for queue rather than adopting the RDMA one ( also why they lack critical feature in term of security and multicore support)
  • ARM Cortex-A72 chips : coming in 2016 thus ARM 64-bit processors that can run at clock speeds of up to 2.5 GHz. They can also be paired with lower-power ARM Cortex-A53
  • Coherency : ARM coherence for heterogeneous core package.
  • Captain Proto : fast data interchange format and capability-based RPC system.


Wednesday, December 03, 2014

Links of the day 3 - 11 - 2014

Today's links 3/12/2014:  #container management , choosing a #nosql platform, future programming , #ARM v8.1 , #Apache drill for #Hadoop

  • Giant Swarm : kubernetes alternative 
  • Visual Guide to NoSQL Systems : picking one is not as trivial as what they want you to believe.
  • Future Programming Workshop 2014 : Video from the workshop 
  • Arm v8 architecture : An overview of the ARMv8.1 architecture extensions (enhancements over current ARMv8). An important feature is the possibility to run a host OS kernel directly in EL2 (Hypervisor mode) saving some extra context switches for KVM.
  • Drill : new top level project from apache , aim to deliver Schema-free SQL Query Engine for Hadoop and NoSQL

Thursday, November 13, 2014

Links of the day 13 - 11 - 2014

Today's links 13/11/2014: Mellanox ConnectX4, Immutable infrastructure and ARM server
  • ConnectX4 : EDR 100Gb/s InfiniBand and 100Gb/s Ethernet, 150M messages/second, impressive numbers from Mellanox. 
  • Fugue: immutable infrastructure delivering Automating the creation and operations of cloud infrastructure, Short-lived and simplified compute instances
  • Custom Cloud Arm Server : online lab design its own ARM based server for its cloud infrastructure.

Tuesday, October 21, 2014

Links of the day 21 - 10 - 2014

Today's links 21/20/2014: all about #Linux #networking with a little bit of  #HPC distributed #storage

  • State of Linux network stack : what's new and interesting in the latest kernel release, especially the low-latency device polling
  • KVM Forum : all videos of this year KVM forum . Some interesting talk especially on the HPC front and an interesting quote from Vincent Jardin: " if you want to have high performance networking or NVF solution don't use virtualization use container"
  • RDMA and ARM : Mellanox bring its RoCE adapter to the moonshot project. Interesting to see what type of application would leverage such architecture combination: a lot of small processors with a fast fabric.  
  • IX : solution that isclose to achieve the holy grail of networking - Low latency with high throughput (line rate)
  • (Fast Forward) Storage and I/O : Distributed Application Object Storage (DAOS) by Intel for HPC solution. A lot of flash , burst buffer with Lustre for supercomputer. Very interesting approach to address the challenge of future exascale computing platform.