Google releases new visual language model PixelLLM

PixeILLM can be applied to a variety of location-aware visual language tasks, including referential localization, positional conditional word mapping, and dense object captioning, and has achieved state-of-the-art performance on RefCOCO and VisualGenome.

Previous:

Next:

Leave a Reply

Please Login to Comment