Fecha de publicación:
--
Fuente:
WIPO "tomato"
Embodiments described herein provide a method of generating a vision-language task output for a text instruction relating to an input image, the method comprising receiving, via a data interface, the input image and the text instruction comprising an instruction relating to the image. The method further includes encoding, via an image encoder, the image into a first image representation. The method further includes adapting, by a multimodal encoder connected to the image encoder, the first image representation to generate a second image representation compatible with a neural network based language model. The method further includes generating, by the neural network based language model connected to the multimodal encoder, the vision-language task output in response to the text instruction based on an input combining the second image representation and the text instruction.