1,720,977 research outputs found
Recommended from our members
Matrix Multiplication is Almost All You Need
Matrix multiplication is a bottleneck operation needed by many classes of algorithms in several domains. As such a ubiquitous operation, many modern computing systems possess specialized hardware accelerators. These accelerators often rely on small, general-purpose host CPUs to orchestrate the entire application and run portions of it that cannot be run on the accelerator incurring expensive communication overhead. In this work, I explore the possibility of instead executing general-purpose code on the matrix-multiplication accelerator to alleviate this overhead and enable applications to be entirely executed within the accelerator. I present a compiler, The Ultimate Mullifier, that targets the ISA of an open source matrix multiplication accelerator, OpenTPU, and takes as its source unmodified programs written in a subset of C. This subset is wide enough to express common algorithms for sorting, searching, and graph traversal with only minor hardware and ISA modifications required. I evaluate these common algorithms as microbenchmarks and show accelerator-only execution can be more efficient in many circumstances
Recommended from our members
Full-System Collaboration in Heterogeneous SoCs with a Hardware Network Stack
As we approach some of the physical limits of transistor scale and performance, modern system-on-chips (SoCs) have become increasingly heterogeneous, incorporating more hardware accelerators in their designs. The general design philosophy has been to treat these accelerators simply as an off-loading co-processor, each with their own custom software drivers. This wastes system performance and kernel developer time and inherently prohibits the amount of collaboration between all of the SoC components. Our project, Pengwing, seeks to alter the current standard surrounding accelerator usage to provide powerful collaboration between hardware and software system services. We implement a novel blended operating system design that follows the idea of software-oriented acceleration. In this paradigm, robust software-hardware interactions are achieved by providing a system substrate that enables communication with accelerators using existing software abstractions like shared-memory queues. My contributions to this project revolve around incorporating Beehive, a hardware network stack, into our SoC design. By intentionally decoupling Beehive from the cores during integration, software and hardware alike can utilize various network functions provided through Beehive using our standard software queue API. This thesis will detail the versatility present in our blended OS implementation and showcase robust software-hardware interaction through various setups that leverage Beehive utilizing the same underlying hardware SoC design
Recommended from our members
Advancing Synthesizable Verilog/SystemVerilog Education with Open-Source Tools and Autograders
In the rapidly expanding semiconductor industry, there is an increasing demand for skilled chip developers. Yet, the steep learning curve associated with Hardware Description Languages (HDLs) often acts as a significant barrier for students hoping to pursue a career in digital design. Drawing upon my experience as a HDL educator, which includes teaching Verilog to UCSB's IEEE student chapter and serving as a Teaching Assistant for UCSB's Verilog courses, I have meticulously developed and refined a comprehensive set of methods and resources for Verilog education. My objective encompassed two key facets: equipping students with quality industry-preparation and kindling passion for exploring hardware design. Through a strategic blend of approaches consisting of the integration of accessible open-source tools, the enforcement of popular coding style guides, the implementation of autograders for personalized feedback, and the incorporation of open-source IP blocks into lessons, students can attain proficiency in designing RTL (Register Transfer Level) for rigorously verified hardware systems. These strategies help reduce Verilog's steep learning curve while also expediting the introduction of more advanced topics in digital design and computer architecture. The methods and resources detailed in this thesis will prepare students for the expectations of the semiconductor industry, enhance their coding skills, and promote an accessible and engaging learning environment, ultimately meeting the growing demand for chip developers
Recommended from our members
Hardware implementation and analysis of memory interfaces to integrate a vector accelerator into a manycore Network-on-Chip
In recent years, there has been a growing demand for vector processors due to their increasing application in deep-learning applications. On the other hand, with the strong need for energy efficiency and high performance, heterogeneous architecture plays an important role and becomes increasingly complex. However, the way of connecting the memory hierarchy to the vector processor in SOC (System-on-Chip) is critical to the system’s performance [8]. This work presents tile design which is based on OpenPiton and BYOC [4] [3]. Tile consists of a 64-bit, single-issue, in-order RISC-V core Ariane [14], along with a 64-bit vector processor ARA [7] [13] which implemented RISC-V V extension version 1.0. This work makes the following contributions. First, it involves the design and implementation of an adapter (bridge) that converts memory request from AMBA AXI to OpenPiton NoC. This adapter enables ARA memory access functionality and facilitates the integration of future accelerators into OpenPiton. Secondly, a tile design is presented, which includes ARA, a RISC-V vector processor, Ariane (a RISC-V core), L1.5 cache, L2 cache, and the implemented bridge. The performance of the tile is evaluated using different versions of bridges connected to the last-level cache (LLC) or off-chip memory. The analysis indicates that a wide data width bridge does not necessarily improve performance significantly. Several factors, such as NoC traffic confliction or unused data fetch, can narrow the performance gap between small and large width bridges. Furthermore, the experiments demonstrate that memory exhibits advantages when dealing with large data widths, and memory saturation also occurs during LLC access. Finally, the thesis proposes the implementation of MSHR (Miss Status Handling Register) and extends this design to manycore architectures to enhance performance
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Recommended from our members
Automated Reasoning for Agile and Robust Chip Design
Modern chip design embodies enormous complexity, from general-purpose processors to specialized hardware accelerators. With the trend towards specialization, chip designers need techniques that let them quickly iterate over a design while fitting into familiar programming languages and tools. However, designing a chip with speed and robustness remains a challenge. Chip design requires reasoning between different layers of abstraction, however these tools do not provide mechanisms to connect specifications with implementations to ensure correctness. Programming languages for chip design rely on technology-specific components, but lack helpful abstractions needed to support common deployment platforms, making it difficult to adapt and compose designs. And further, the design ecosystem is fragmented between systems and tool chains without the ability to interoperate.This thesis presents my research on improving chip design tools with automated reasoning techniques. I use program synthesis techniques to bridge the gap between an architectural specification and a low-level hardware implementation, developing a new technique called control logic synthesis. I establish a new field called hardware decompilation, which is about lifting common hardware artifacts to high-level source code, enabling design transpilation and automating the effort of re-targeting designs to different technologies. And finally, to address challenges with technology constraints, I developed a memory design language that uses equational reasoning techniques to automatically target multiple memory technologies from a single interface. Through the application of these automated reasoning techniques, I opened two wholly new areas in the chip design space enabling novel design processes that were not possible before, improving developer agility and design verifiability
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
