A Manycore Architecture for Scalability, Programmability, and Compute Density

dc.contributor.advisorTaylor, Michael
dc.contributor.authorJung, Dai Cheol
dc.date.accessioned2026-08-11T19:28:35Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractManycore architectures have been proposed as a way to transform abundant silicon resources into highly programmable parallel processors with exceptional compute density. However, many architectural mechanisms inherited from large, complex cores – designed to optimize single-thread performance and support cache coherence – have limited their ability to achieve the compute density and scalability needed to compete with other parallel architectures. This thesis presents HammerBlade Manycore, which integrates numerous novel ideas spanning multiple areas of parallel architecture, such as VLSI resource organization, on-chip networks, memory hierarchy, memory-level parallelism, address mapping and translation, synchronization, and thread management. Its flexibility has been demonstrated with a parallel benchmark suite representing a broad range of computational and communication patterns. Network-on-Chip (NoC) has become one of the most critical components in modern, parallel architectures. This thesis presents Ruche Networks, which leverage unused wiring resources cost-effectively to reduce network diameter and increase bisection bandwidth by augmenting the 2-D mesh with uniform long-range physical links. Ruche Networks are tileable, physically scalable, and energy efficient. A comprehensive evaluation is provided, comparing Ruche Networks with other NoCs, such as 2-D mesh, torus, and multi-mesh, in terms of area, power, network performance, and scalability. Finally, this thesis presents the 12 nm implementation of the 2048-core HammerBlade ASIC. This chip achieves several milestones: it is the first to implement Ruche Networks in a large-scale and high-frequency manycore processor, and it sets a record peak RISC-V instruction throughput and CoreMark scores. Building a chip at this scale presents significant challenges, including timing closure and excessively long CAD tool runtimes. This thesis documents the technical lessons learned during tapeout.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherJung_washington_0250E_29430.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57315
dc.language.isoen_US
dc.rightsnone
dc.subjectComputer engineering
dc.subject.otherElectrical and computer engineering
dc.titleA Manycore Architecture for Scalability, Programmability, and Compute Density
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Jung_washington_0250E_29430.pdf
Size:
5.4 MB
Format:
Adobe Portable Document Format