To interconnect their growing number of servers, current supercomputers and data centers are starting to adopt low-diameter networks, such as HyperX, Dragonfly and Dragonfly+. These emergent topologies require balancing the load over their links and finding suitable non-minimal routing mechanisms for them becomes particularly challenging. The Valiant load balancing scheme is a very popular choice for non-minimal routing. Evolved adaptive routing mechanisms implemented in real systems are based on this Valiant scheme. All these low-diameter networks are deadlock-prone when non-minimal routing is employed. Routing deadlocks occur when packets cannot progress due to cyclic dependencies. Therefore, developing efficient deadlock-free packet routing mechanisms is critical for the progress of these emergent networks. The routing function includes the routing algorithm for path selection and the buffers management policy that dictates how packets allocate the buffers of the switches on their paths. For the same routing algorithm, a different buffer management mechanism can lead to a very different performance. Moreover, certain mechanisms considered efficient for avoiding deadlocks, may still suffer from hard to pinpoint instabilities that make erratic the network response. This paper focuses on exploring the impact of these buffers management policies on the performance of current interconnection networks, showing a 90\% of performance drop if an incorrect buffers management policy is used. Moreover, this study not only characterizes some of these undesirable scenarios but also proposes practicable solutions.
翻译:为了互联日益增长的服务器数量,当前的超级计算机和数据中心开始采用低直径网络,例如HyperX、Dragonfly和Dragonfly+。这些新兴的拓扑结构需要在其链路上均衡负载,而为其寻找合适的非最小路由机制变得尤为具有挑战性。Valiant负载均衡方案是非最小路由中非常流行的一种选择。在实际系统中实现的演进式自适应路由机制均基于此Valiant方案。当采用非最小路由时,所有这些低直径网络都容易出现死锁。当报文因循环依赖而无法前进时,就会发生路由死锁。因此,为这些新兴网络开发高效的无死锁报文路由机制至关重要。路由功能包括用于路径选择的路由算法,以及决定报文如何在其路径上分配交换机缓冲区的缓冲管理策略。对于相同的路由算法,不同的缓冲管理机制可能导致截然不同的性能。此外,某些被认为能有效避免死锁的机制,可能仍然存在难以定位的不稳定性,从而使网络响应变得不稳定。本文重点探讨这些缓冲管理策略对当前互连网络性能的影响,结果表明,如果使用不正确的缓冲管理策略,性能会下降90%。此外,这项研究不仅描述了其中一些不良场景的特征,还提出了可行的解决方案。