A Manycore Architecture for Scalability, Programmability, and Compute Density

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Manycore architectures have been proposed as a way to transform abundant silicon resources into highly programmable parallel processors with exceptional compute density. However, many architectural mechanisms inherited from large, complex cores – designed to optimize single-thread performance and support cache coherence – have limited their ability to achieve the compute density and scalability needed to compete with other parallel architectures. This thesis presents HammerBlade Manycore, which integrates numerous novel ideas spanning multiple areas of parallel architecture, such as VLSI resource organization, on-chip networks, memory hierarchy, memory-level parallelism, address mapping and translation, synchronization, and thread management. Its flexibility has been demonstrated with a parallel benchmark suite representing a broad range of computational and communication patterns. Network-on-Chip (NoC) has become one of the most critical components in modern, parallel architectures. This thesis presents Ruche Networks, which leverage unused wiring resources cost-effectively to reduce network diameter and increase bisection bandwidth by augmenting the 2-D mesh with uniform long-range physical links. Ruche Networks are tileable, physically scalable, and energy efficient. A comprehensive evaluation is provided, comparing Ruche Networks with other NoCs, such as 2-D mesh, torus, and multi-mesh, in terms of area, power, network performance, and scalability. Finally, this thesis presents the 12 nm implementation of the 2048-core HammerBlade ASIC. This chip achieves several milestones: it is the first to implement Ruche Networks in a large-scale and high-frequency manycore processor, and it sets a record peak RISC-V instruction throughput and CoreMark scores. Building a chip at this scale presents significant challenges, including timing closure and excessively long CAD tool runtimes. This thesis documents the technical lessons learned during tapeout.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI