Tsinghua Joint Byte Open Source Auditory Big Language Model SALMONN

Tsinghua University, in collaboration with the Byte Jump Volcano Speech Team, has launched SALMONN, a cognitively oriented open-source auditory macrolanguage model, which is not only capable of perceiving and understanding various types of audio inputs, and can perform English speech recognition, speech translation, emotion recognition, and audio subtitle generation tasks, but also springs up with a wide range of unlearned multilingual and cross-modal capabilities.

Previous:

Next:

Leave a Reply

Please Login to Comment