The text generation process of large language models (LLMs) is expensive and slow, reducing inference speed. The University of Waterloo, the Canadian Vector Institute, and Peking University have jointly released EAGLE, which aims to improve the inference speed of large language models while ensuring a consistent distribution of model output text.
Large Model Reasoning Efficiency Increases 3X Without Loss, University of Waterloo, Peking University and Other Institutions Release EAGLE
Previous: 加速算力基础设施建设,「数智说」算力新基建论坛即将启幕