Even watching Hollywood blockbusters has learned! Jiajia's team uses 2token to make big models roll to a new level

LLaMA-VID, a multimodal macromodel created by Jiajia's team at the Chinese University of Hong Kong, can handle inputs from movies up to 3 hours long, which fills a gap in the field of macromodels for long videos. The distinguishing feature of this system is its ability to encode each video frame in the form of two tokens, a simple yet effective approach. This model shows great accuracy in understanding game trailers, analyzing short videos and images, and elaborating movie plots. The power of its performance is highly appreciated by researchers.

Previous:

Next:

Leave a Reply

Please Login to Comment