-
公开(公告)号:US20240244287A1
公开(公告)日:2024-07-18
申请号:US18154412
申请日:2023-01-13
Applicant: Adobe Inc.
Inventor: Kim Pascal PIMMEL , Stephen Joseph DIVERDI , Jiaju MA , Rubaiat HABIB , Li-Yi WEI , Hijung SHIN , Deepali ANEJA , John G. NELSON , Wilmot LI , Dingzeyu LI , Lubomira Assenova DONTCHEVA , Joel Richard BRANDT
IPC: H04N21/431 , G06F3/04812 , G06F3/0482 , H04N21/4402
CPC classification number: H04N21/4312 , G06F3/04812 , G06F3/0482 , H04N21/440236
Abstract: Embodiments of the present disclosure provide, a method, a system, and a computer storage media that provide mechanisms for multimedia effect addition and editing support for text-based video editing tools. The method includes generating a user interface (UI) displaying a transcript of an audio track of a video and receiving, via the UI, input identifying selection of a text segment from the transcript. The method also includes in response to receiving, via the UI, input identifying selection of a particular type of text stylization or layout for application to the text segment. The method further includes identifying a video effect corresponding to the particular type of text stylization or layout, applying the video effect to a video segment corresponding to the text segment, and applying the particular type of text stylization or layout to the text segment to visually represent the video effect in the transcript.
-
公开(公告)号:US20250139161A1
公开(公告)日:2025-05-01
申请号:US18431134
申请日:2024-02-02
Applicant: ADOBE INC.
Inventor: Deepali ANEJA , Zeyu JIN , Hijung SHIN , Anh Lan TRUONG , Dingzeyu LI , Hanieh DEILAMSALEHY , Rubaiat HABIB , Matthew David FISHER , Kim Pascal PIMMEL , Wilmot LI , Lubomira Assenova DONTCHEVA
IPC: G06F16/783 , G06F16/738 , G06V20/40 , G06V40/16
Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for cutting down a user's larger input video into an edited video comprising the most important video segments and applying corresponding video effects. Some embodiments of the present invention are directed to adding captioning video effects to the trimmed video (e.g., applying face-aware and non-face-aware captioning to emphasize extracted video segment headings, important sentences, quotes, words of interest, extracted lists, etc.). For example, a prompt is provided to a generative language model to identify portions of a transcript (e.g., extracted scene summaries, important sentences, lists of items discussed in the video, etc.) to apply to corresponding video segments as captions depending on the type of caption (e.g., an extracted heading may be captioned at the start of a corresponding video segment, important sentences and/or extracted list items may be captioned when they are spoken).
-
公开(公告)号:US20250168442A1
公开(公告)日:2025-05-22
申请号:US19033062
申请日:2025-01-21
Applicant: Adobe Inc.
Inventor: Kim Pascal PIMMEL , Stephen Joseph DIVERDI , Jiaju MA , Rubaiat HABIB , LI-Yi WEI , Hijung SHIN , Deepali ANEJA , John G. NELSON , Wilmot LI , Dingzeyu LI , Lubomira Assenova DONTCHEVA , Joel Richard BRANDT
IPC: H04N21/431 , G06F3/04812 , G06F3/0482 , H04N21/4402
Abstract: Embodiments of the present disclosure provide, a method, a system, and a computer storage media that provide mechanisms for multimedia effect addition and editing support for text-based video editing tools. The method includes generating a user interface (UI) displaying a transcript of an audio track of a video and receiving, via the UI, input identifying selection of a text segment from the transcript. The method also includes in response to receiving, via the UI, input identifying selection of a particular type of text stylization or layout for application to the text segment. The method further includes identifying a video effect corresponding to the particular type of text stylization or layout, applying the video effect to a video segment corresponding to the text segment, and applying the particular type of text stylization or layout to the text segment to visually represent the video effect in the transcript.
-
公开(公告)号:US20240134909A1
公开(公告)日:2024-04-25
申请号:US17967703
申请日:2022-10-17
Applicant: Adobe Inc.
Inventor: Lubomira Assenova DONTCHEVA , Dingzeyu LI , Kim Pascal PIMMEL , Hijung SHIN , Hanieh DEILAMSALEHY , Aseem Omprakash AGARWALA , Joy Oakyung KIM , Joel Richard BRANDT , Cristin Ailidh Fraser
IPC: G06F16/732
CPC classification number: G06F16/732
Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for a visual and text search interface used to navigate a video transcript. In an example embodiment, a freeform text query triggers a visual search for frames of a loaded video that match the freeform text query (e.g., frame embeddings that match a corresponding embedding of the freeform query), and triggers a text search for matching words from a corresponding transcript or from tags of detected features from the loaded video. Visual search results are displayed (e.g., in a row of tiles that can be scrolled to the left and right), and textual search results are displayed (e.g., in a row of tiles that can be scrolled up and down). Selecting (e.g., clicking or tapping on) a search result tile navigates a transcript interface to a corresponding portion of the transcript.
-
公开(公告)号:US20240134597A1
公开(公告)日:2024-04-25
申请号:US17967714
申请日:2022-10-17
Applicant: Adobe Inc.
Inventor: Lubomira Assenova DONTCHEVA , Anh Lan TRUONG , Hanieh DEILAMSALEHY , Kim Pascal PIMMEL , Aseem Omprakash AGARWALA , Dingzeyu Li , Joel Richard BRANDT , Joy Oakyung KIM
IPC: G06F3/16 , G06F3/0482 , G06F3/0484 , G06F16/735 , G06F16/738
CPC classification number: G06F3/167 , G06F3/0482 , G06F3/0484 , G06F16/735 , G06F16/738
Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for a question search for meaningful questions that appear in a video. In an example embodiment, an audio track from a video is transcribed, and the transcript is parsed to identify sentences that end with a question mark. Depending on the embodiment, one or more types of questions are filtered out, such as short questions less than a designated length or duration, logistical questions, and/or rhetorical questions. As such, in response to a command to perform a question search, the questions are identified, and search result tiles representing video segments of the questions are presented. Selecting (e.g., clicking or tapping on) a search result tile navigates a transcript interface to a corresponding portion of the transcript.
-
公开(公告)号:US20250140292A1
公开(公告)日:2025-05-01
申请号:US18431103
申请日:2024-02-02
Applicant: ADOBE INC.
Inventor: Anh Lan TRUONG , Deepali ANEJA , Hijung SHIN , Rubaiat HABIB , Jakub FISER , Kishore RADHAKRISHNA , Joel Richard BRANDT , Matthew David FISHER , Zeyu JIN , Kim Pascal PIMMEL , Wilmot LI , Lubomira Assenova DONTCHEVA
IPC: G11B27/036 , G06V20/40 , G06V40/16 , H04N5/262
Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for cutting down a user's larger input video into an edited video comprising the most important video segments and applying corresponding video effects. Some embodiments of the present invention are directed to adding face-aware scale magnification to the trimmed video (e.g., applying scale magnification to simulate a camera zoom effect that hides shot cuts with respect to the subject's face). For example, as the trimmed video transitions from one video segment to the next video segment, a scale magnification may be applied that zooms in on a detected face at a boundary between the video segments to smooth the transition between video segments.
-
公开(公告)号:US20250140291A1
公开(公告)日:2025-05-01
申请号:US18431139
申请日:2024-02-02
Applicant: ADOBE INC.
Inventor: Hanieh DEILAMSALEHY , Jui-Hsien WANG , Zhengyang MA , Dingzeyu LI , Hijung SHIN , Aseem Omprakash AGARWALA , Kim Pascal PIMMEL , Lubomira Assenova DONTCHEVA
IPC: G11B27/031 , G06V20/40 , G10L15/04 , G10L15/183 , G10L21/0272 , G10L25/57 , G11B27/06 , G11B27/34
Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for identifying the relevant segments that effectively summarize the larger input video and/or form a rough cut, and assembling them into one or more smaller trimmed videos. For example, visual scenes and corresponding scene captions are extracted from the input video and associated with an extracted diarized and timestamped transcript to generate an augmented transcript. The augmented transcript is applied to a large language model to extract sentences that characterize a trimmed version of the input video (e.g., a natural language summary, a representation of identified sentences from the transcript). As such, corresponding video segments are identified (e.g., using similarity to match each sentence in a generated summary with a corresponding transcript sentence) and assembled into one or more trimmed videos. In some embodiments, the trimmed video is generated based on a user's query and/or desired length.
-
-
-
-
-
-