Mesh-of-Trees and Alternative Interconnection Networks for Single-Chip Parallelism

TitleMesh-of-Trees and Alternative Interconnection Networks for Single-Chip Parallelism
Publication TypeJournal Articles
Year of Publication2009
AuthorsBalkan AO, Qu G, Vishkin U
JournalVery Large Scale Integration (VLSI) Systems, IEEE Transactions on
Pagination1419 - 1432
Date Published2009/10//
ISBN Number1063-8210
Keywords90, cache;single, complexity;multiprocessor, delay;single-chip, first-level, high-throughput, interconnection, low-latency, network;memory, network;shared, networks;network-on-chip;parallel, nm;wire, Parallel, parallelism;size, processing;, processor;single-chip, switch, topologies;on-chip, units;mesh-of-trees;network

In single-chip parallel processors, it is crucial to implement a high-throughput low-latency interconnection network to connect the on-chip components, especially the processing units and the memory units. In this paper, we propose a new mesh of trees (MoT) implementation of the interconnection network and evaluate it relative to metrics such as wire complexity, total register count, single switch delay, maximum throughput, tradeoffs between throughput and latency, and post-layout performance. We show that on-chip interconnection networks can provide higher bandwidth between processors and shared first-level cache than previously considered possible, facilitating greater scalability of memory architectures that require that. MoT is also compared, both analytically and experimentally, to some other traditional network topologies, such as hypercube, butterfly, fat trees and butterfly fat trees. When we evaluate a 64-terminal MoT network at 90-nm technology, concrete results show that MoT provides higher throughput and lower latency especially when the input traffic (or the on-chip parallelism) is high, at comparable area. A recurring problem in networking and communication is that of achieving good sustained throughput in contrast to just high theoretical peak performance that does not materialize for typical work loads. Our quantitative results demonstrate a clear advantage of the proposed MoT network in the context of single-chip parallel processing.