TRANSCRIBING PHOTOGRAPHIC IMAGES

Fuente: WIPO "tomato"
A method can include obtaining an image of structured text. Transcription text can be generated based on the image using a transcription model. A set of features can be determined based on the image. A first feature subset can be extracted from the image. A second feature subset can be extracted from the transcription text. A first image model, including a first image convolutional network or a first transformer model, can generate an image representation vector using the first feature subset. A tabular layer model can generate a transcription representation vector using the second feature subset. A combination vector can be generated by combining the image and transcription representation vectors. A machine learning classification model can determine an accuracy classification for the transcription text using the combination vector. Responsive to the accuracy classification indicating that the transcription text is accurate, the transcription text can be displayed within an application.