Search Results for author: Mengjia Dai

Found 1 papers, 1 papers with code

Allo: A Programming Model for Composable Accelerator Design

2 code implementations7 Apr 2024 Hongzheng Chen, Niansong Zhang, Shaojie Xiang, Zhichen Zeng, Mengjia Dai, Zhiru Zhang

For the GPT2 model, the inference latency of the Allo generated accelerator is 1. 7x faster than the NVIDIA A100 GPU with 5. 4x higher energy efficiency, demonstrating the capability of Allo to handle large-scale designs.

Cannot find the paper you are looking for? You can Submit a new open access paper.