Patent search ap:("Adobe Inc.") AND inv:"Fabian David CABA HEILBRON" Page 1

1.

发明公开
FACE-AWARE SPEAKER DIARIZATION FOR TRANSCRIPTS AND TEXT-BASED VIDEO EDITING 审中-公开

公开(公告)号：US20240127857A1

公开(公告)日：2024-04-18

申请号：US17967399

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Fabian David CABA HEILBRON , Xue BAI , Aseem Omprakash AGARWALA , Haoran CAI , Lubomira Assenova DONTCHEVA

IPC: G11B27/031 , G06V20/40

CPC classification number: G11B27/031 , G06V20/41

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for face-aware speaker diarization. In an example embodiment, an audio-only speaker diarization technique is applied to generate an audio-only speaker diarization of a video, an audio-visual speaker diarization technique is applied to generate a face-aware speaker diarization of the video, and the audio-only speaker diarization is refined using the face-aware speaker diarization to generate a hybrid speaker diarization that links detected faces to detected voices. In some embodiments, to accommodate videos with small faces that appear pixelated, a cropped image of any given face is extracted from each frame of the video, and the size of the cropped image is used to select a corresponding active speaker detection model to predict an active speaker score for the face in the cropped image.

2.

发明公开
MUSIC-AWARE SPEAKER DIARIZATION FOR TRANSCRIPTS AND TEXT-BASED VIDEO EDITING 审中-公开

公开(公告)号：US20240127820A1

公开(公告)日：2024-04-18

申请号：US17967502

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Justin Jonathan SALAMON , Fabian David CABA HEILBRON , Xue BAI , Aseem Omprakash AGARWALA , Hijung SHIN , Lubomira Assenova DONTCHEVA

IPC: G10L15/26 , G11B27/031

CPC classification number: G10L15/26 , G11B27/031

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for music-aware speaker diarization. In an example embodiment, one or more audio classifiers detect speech and music independently of each other, which facilitates detecting regions in an audio track that contain music but do not contain speech. These music-only regions are compared to the transcript, and any transcription and speakers that overlap in time with the music-only regions are removed from the transcript. In some embodiments, rather than having the transcript display the text from this detected music, a visual representation of the audio waveform is included in the corresponding regions of the transcript.

Patent Agency Ranking