METHOD OF PROCESSING MULTIMEDIA DATA, DEVICE AND MEDIUM

    公开(公告)号:US20230115737A1

    公开(公告)日:2023-04-13

    申请号:US18080432

    申请日:2022-12-13

    Abstract: A method of processing multimedia data, a device, and a medium, which relates to a field of an artificial intelligence technology, in particular to fields of knowledge graph and deep learning. The method of processing the multimedia data includes: recognizing the multimedia data so as to obtain at least one key information of the multimedia data; querying a predetermined knowledge base according to the at least one key information, so as to determine a multimedia name associated with the at least one key information and an association degree between the multimedia name and the at least one key information; and determining, in the multimedia name, a name of the multimedia data based on a similarity between alternative multimedia data for the multimedia name and the multimedia data, in response to the association degree being less than a first threshold value.

    MULTIMODAL DATA PROCESSING
    3.
    发明申请

    公开(公告)号:US20230010160A1

    公开(公告)日:2023-01-12

    申请号:US17945415

    申请日:2022-09-15

    Abstract: Disclosed are a method for processing multimodal data using a neural network, a device, and a medium, and relates to the field of artificial intelligence and, in particular to multimodal data processing, video classification, and deep learning. The neural network includes: an input subnetwork configured to receive the multimodal data to output respective first features of a plurality of modalities; a plurality of cross-modal feature subnetworks, each of which is configured to receive respective first features of two corresponding modalities to output a cross-modal feature corresponding to the two modalities; a plurality of cross-modal fusion subnetworks, each of which is configured to receive at least one cross-modal feature corresponding to a corresponding target modality and other modalities to output a second feature of the target modality; and an output subnetwork configured to receive respective second features of the plurality of modalities to output a processing result of the multimodal data.

    QUESTION ANSWERING METHOD, METHOD OF TRAINING A QUESTION ANSWERING MODEL, ELECTRONIC DEVICE, AND MEDIUM

    公开(公告)号:US20230153337A1

    公开(公告)日:2023-05-18

    申请号:US18157452

    申请日:2023-01-20

    CPC classification number: G06F16/3329 G06F40/30

    Abstract: A question answering method, a method of training a question answering model, a device, and a medium are provided, which relate to a field of artificial intelligence technology, in particular to fields of natural language processing technology, deep learning technology, and knowledge mapping technology. The question answering method includes: obtaining data to be processed, wherein the data to be processed includes a question and candidate answers; performing general semantic understanding on the data to be processed to obtain a general data feature; selecting a target question answering mode from candidate question answering modes based on the general data feature; and processing the general data feature by using the target question answering mode, to obtain a target answer for the question from the candidate answers.

    INFORMATION SEARCH METHOD AND DEVICE, ELECTRONIC DEVICE, AND STORAGE MEDIUM

    公开(公告)号:US20230008897A1

    公开(公告)日:2023-01-12

    申请号:US17932598

    申请日:2022-09-15

    Abstract: An information search method includes: obtaining search words at least including a question to be searched and obtaining an initial text vector representation of the search words; obtaining a video corresponding to the search words, and obtaining multi-modality vector representations of the video; starting from the initial text vector representation, performing N rounds of interaction between the video and the search words based on the multi-modality vector representations and a text vector representation of the search words of a current round, to generate a target fusion vector representation, where N is an integer greater than or equal to 1; and obtaining target video frames matching the question to be searched by annotating the video based on the target fusion vector representation.

Patent Agency Ranking