Patent search ap:("Adobe Inc.") AND inv:"Aseem Omprakash AGARWALA" Page 1

1.

发明公开
TRANSCRIPT PARAGRAPH SEGMENTATION AND VISUALIZATION OF TRANSCRIPT PARAGRAPHS 审中-公开

公开(公告)号：US20240126994A1

公开(公告)日：2024-04-18

申请号：US17967562

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Hanieh DEILAMSALEHY , Aseem Omprakash AGARWALA , Haoran CAI , Hijung SHIN , Joel Richard BRANDT , Lubomira Assenova DONTCHEVA

IPC: G06F40/30 , G06F40/205 , H04N5/93

CPC classification number: G06F40/30 , G06F40/205 , H04N5/9305

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for segmenting a transcript into paragraphs. In an example embodiment, a transcript is segmented to start a new paragraph whenever there is a change in speaker and/or a long pause in speech. If any remaining paragraphs are longer than a designated length or duration (e.g., 50 or 100 words), each of those paragraphs is segmented using dynamic programming to minimize a cost function that penalizes candidate paragraphs based on divergence from a target paragraph length and/or that rewards candidate paragraphs that group semantically similar sentences. As such, the transcript is visualized, segmented at the identified paragraphs.

2.

发明公开
VIDEO SEGMENT SELECTION AND EDITING USING TRANSCRIPT INTERACTIONS 审中-公开

公开(公告)号：US20240135973A1

公开(公告)日：2024-04-25

申请号：US17967364

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Xue BAI , Justin Jonathan SALAMON , Aseem Omprakash AGARWALA , Hijung SHIN , Haoran CAI , Joel Richard BRANDT , Lubomira Assenova DONTCHEVA , Cristin Ailidh Fraser

IPC: G11B27/036 , G06F40/166 , G10L15/26 , G10L25/57 , G11B27/34

CPC classification number: G11B27/036 , G06F40/166 , G10L15/26 , G10L25/57 , G11B27/34 , G06F3/0482

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for identifying candidate boundaries for video segments, video segment selection using those boundaries, and text-based video editing of video segments selected via transcript interactions. In an example implementation, boundaries of detected sentences and words are extracted from a transcript, the boundaries are retimed into an adjacent speech gap to a location where voice or audio activity is a minimum, and the resulting boundaries are stored as candidate boundaries for video segments. As such, a transcript interface presents the transcript, interprets input selecting transcript text as an instruction to select a video segment with corresponding boundaries selected from the candidate boundaries, and interprets commands that are traditionally thought of as text-based operations (e.g., cut, copy, paste) as an instruction to perform a corresponding video editing operation using the selected video segment.

3.

发明公开
FACE-AWARE SPEAKER DIARIZATION FOR TRANSCRIPTS AND TEXT-BASED VIDEO EDITING 审中-公开

公开(公告)号：US20240127857A1

公开(公告)日：2024-04-18

申请号：US17967399

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Fabian David CABA HEILBRON , Xue BAI , Aseem Omprakash AGARWALA , Haoran CAI , Lubomira Assenova DONTCHEVA

IPC: G11B27/031 , G06V20/40

CPC classification number: G11B27/031 , G06V20/41

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for face-aware speaker diarization. In an example embodiment, an audio-only speaker diarization technique is applied to generate an audio-only speaker diarization of a video, an audio-visual speaker diarization technique is applied to generate a face-aware speaker diarization of the video, and the audio-only speaker diarization is refined using the face-aware speaker diarization to generate a hybrid speaker diarization that links detected faces to detected voices. In some embodiments, to accommodate videos with small faces that appear pixelated, a cropped image of any given face is extracted from each frame of the video, and the size of the cropped image is used to select a corresponding active speaker detection model to predict an active speaker score for the face in the cropped image.

4.

发明公开
VISUAL AND TEXT SEARCH INTERFACE FOR TEXT-BASED VIDEO EDITING 审中-公开

公开(公告)号：US20240134909A1

公开(公告)日：2024-04-25

申请号：US17967703

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Lubomira Assenova DONTCHEVA , Dingzeyu LI , Kim Pascal PIMMEL , Hijung SHIN , Hanieh DEILAMSALEHY , Aseem Omprakash AGARWALA , Joy Oakyung KIM , Joel Richard BRANDT , Cristin Ailidh Fraser

IPC: G06F16/732

CPC classification number: G06F16/732

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for a visual and text search interface used to navigate a video transcript. In an example embodiment, a freeform text query triggers a visual search for frames of a loaded video that match the freeform text query (e.g., frame embeddings that match a corresponding embedding of the freeform query), and triggers a text search for matching words from a corresponding transcript or from tags of detected features from the loaded video. Visual search results are displayed (e.g., in a row of tiles that can be scrolled to the left and right), and textual search results are displayed (e.g., in a row of tiles that can be scrolled up and down). Selecting (e.g., clicking or tapping on) a search result tile navigates a transcript interface to a corresponding portion of the transcript.

5.

发明公开
TRANSCRIPT QUESTION SEARCH FOR TEXT-BASED VIDEO EDITING 审中-公开

公开(公告)号：US20240134597A1

公开(公告)日：2024-04-25

申请号：US17967714

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Lubomira Assenova DONTCHEVA , Anh Lan TRUONG , Hanieh DEILAMSALEHY , Kim Pascal PIMMEL , Aseem Omprakash AGARWALA , Dingzeyu Li , Joel Richard BRANDT , Joy Oakyung KIM

IPC: G06F3/16 , G06F3/0482 , G06F3/0484 , G06F16/735 , G06F16/738

CPC classification number: G06F3/167 , G06F3/0482 , G06F3/0484 , G06F16/735 , G06F16/738

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for a question search for meaningful questions that appear in a video. In an example embodiment, an audio track from a video is transcribed, and the transcript is parsed to identify sentences that end with a question mark. Depending on the embodiment, one or more types of questions are filtered out, such as short questions less than a designated length or duration, logistical questions, and/or rhetorical questions. As such, in response to a command to perform a question search, the questions are identified, and search result tiles representing video segments of the questions are presented. Selecting (e.g., clicking or tapping on) a search result tile navigates a transcript interface to a corresponding portion of the transcript.

6.

发明申请
VIDEO ASSEMBLY USING GENERATIVE ARTIFICIAL INTELLIGENCE 有权

公开(公告)号：US20250140291A1

公开(公告)日：2025-05-01

申请号：US18431139

申请日：2024-02-02

Applicant: ADOBE INC.

Inventor： Hanieh DEILAMSALEHY , Jui-Hsien WANG , Zhengyang MA , Dingzeyu LI , Hijung SHIN , Aseem Omprakash AGARWALA , Kim Pascal PIMMEL , Lubomira Assenova DONTCHEVA

IPC: G11B27/031 , G06V20/40 , G10L15/04 , G10L15/183 , G10L21/0272 , G10L25/57 , G11B27/06 , G11B27/34

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for identifying the relevant segments that effectively summarize the larger input video and/or form a rough cut, and assembling them into one or more smaller trimmed videos. For example, visual scenes and corresponding scene captions are extracted from the input video and associated with an extracted diarized and timestamped transcript to generate an augmented transcript. The augmented transcript is applied to a large language model to extract sentences that characterize a trimmed version of the input video (e.g., a natural language summary, a representation of identified sentences from the transcript). As such, corresponding video segments are identified (e.g., using similarity to match each sentence in a generated summary with a corresponding transcript sentence) and assembled into one or more trimmed videos. In some embodiments, the trimmed video is generated based on a user's query and/or desired length.

7.

发明公开
SPEAKER THUMBNAIL SELECTION AND SPEAKER VISUALIZATION IN DIARIZED TRANSCRIPTS FOR TEXT-BASED VIDEO 审中-公开

公开(公告)号：US20240127855A1

公开(公告)日：2024-04-18

申请号：US17967697

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Lubomira Assenova DONTCHEVA , Xue BAI , Aseem Omprakash AGARWALA , Joel Richard BRANDT

IPC: G11B27/02 , G06V20/40 , G06V40/16

CPC classification number: G11B27/02 , G06V20/46 , G06V20/48 , G06V20/49 , G06V40/172 , G06V40/176

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for selection of the best image of a particular speaker's face in a video, and visualization in a diarized transcript. In an example embodiment, candidate images of a face of a detected speaker are extracted from frames of a video identified by a detected face track for the face, and a representative image of the detected speaker's face is selected from the candidate images based on image quality, facial emotion (e.g., using an emotion classifier that generates a happiness score), a size factor (e.g., favoring larger images), and/or penalizing images that appear towards the beginning or end of a face track. As such, each segment of the transcript is presented with the representative image of the speaker who spoke that segment and/or input is accepted changing the representative image associated with each speaker.

8.

发明公开
MUSIC-AWARE SPEAKER DIARIZATION FOR TRANSCRIPTS AND TEXT-BASED VIDEO EDITING 审中-公开

公开(公告)号：US20240127820A1

公开(公告)日：2024-04-18

申请号：US17967502

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Justin Jonathan SALAMON , Fabian David CABA HEILBRON , Xue BAI , Aseem Omprakash AGARWALA , Hijung SHIN , Lubomira Assenova DONTCHEVA

IPC: G10L15/26 , G11B27/031

CPC classification number: G10L15/26 , G11B27/031

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for music-aware speaker diarization. In an example embodiment, one or more audio classifiers detect speech and music independently of each other, which facilitates detecting regions in an audio track that contain music but do not contain speech. These music-only regions are compared to the transcript, and any transcription and speakers that overlap in time with the music-only regions are removed from the transcript. In some embodiments, rather than having the transcript display the text from this detected music, a visual representation of the audio waveform is included in the corresponding regions of the transcript.

Search Results

Country/Region

Patent validity

Application date

Publication (announcement) day

applicant

The country/region where the applicant is located

Inventor

IPC

IPC Department

IPC class

IPC subclass

IPC group

IPC team

Appearance classification