The Little Alpaca team has developed a decoding algorithm that allows the model to predict the number of 100 tokens 1.5-2.3 times faster, thus accelerating LLM inference. It primarily utilizes the Jacobi iterative method to break the order dependency in autoregressive decoding for the first time.
Transformer's new decoding algorithm is on fire, from the Little Alpaca team | code is open source!
Previous: 美国降息和英伟达财报推动投资者买入AI基金
Next: 字节跳动成立新部门Flow 发力AI应用层