专利检索 ap:("SoundHound, Inc.") AND inv:"Ethan Coeytaux" 第 1 页

1.

发明授权
Method and system for conversation transcription with metadata 有权

公开(公告)号：US12125487B2

公开(公告)日：2024-10-22

申请号：US17450551

申请日：2021-10-11

申请人： SoundHound, Inc.

发明人： Kiersten L. Bradley , Ethan Coeytaux , Ziming Yin

IPC分类号： G10L15/26 , G06F40/134 , G06F40/166 , G06F40/284 , G10L15/02 , G10L15/06 , G10L15/07

CPC分类号： G10L15/26 , G06F40/134 , G06F40/166 , G06F40/284 , G10L15/02 , G10L15/063 , G10L15/07 , G10L2015/0631

摘要： Methods and systems for enabling an efficient review of meeting content via a metadata-enriched, speaker-attributed and multiuser-editable transcript are disclosed. By incorporating speaker diarization and other metadata, the system can provide a structured and effective way to review and/or edit the transcript by one or more editors. One type of metadata can be image or video data to represent the meeting content. Furthermore, the present subject matter utilizes a multimodal diarization model to identify and label different speakers. The system can synchronize various sources of data, e.g., audio channel data, voice feature vectors, acoustic beamforming, image identification, and extrinsic data, to implement speaker diarization.

2.

发明授权
Method and system for conversation transcription with metadata 有权

公开(公告)号：US12020708B2

公开(公告)日：2024-06-25

申请号：US17450552

申请日：2021-10-11

申请人： SoundHound, Inc.

发明人： Kiersten L. Bradley , Ethan Coeytaux , Ziming Yin

IPC分类号： G10L15/26 , G06F40/134 , G06F40/166 , G06F40/284 , G10L15/02 , G10L15/06 , G10L15/07

CPC分类号： G10L15/26 , G06F40/134 , G06F40/166 , G06F40/284 , G10L15/02 , G10L15/063 , G10L15/07 , G10L2015/0631

摘要： Methods and systems for enabling an efficient review of meeting content via a metadata-enriched, speaker-attributed transcript are disclosed. By incorporating speaker diarization and other metadata, the system can provide a structured and effective way to review and/or edit the transcript. One type of metadata can be image or video data to represent the meeting content. Furthermore, the present subject matter utilizes a multimodal diarization model to identify and label different speakers. The system can synchronize various sources of data, e.g., audio channel data, voice feature vectors, acoustic beamforming, image identification, and extrinsic data, to implement speaker diarization.

3.

发明申请
VIDEO CONFERENCE CAPTIONING 有权

公开(公告)号：US20210074298A1

公开(公告)日：2021-03-11

申请号：US16567760

申请日：2019-09-11

申请人： SoundHound, Inc.

发明人： Ethan Coeytaux

IPC分类号： G10L15/26 , G10L15/02 , G10L15/14 , G10L15/19 , G10L19/005

摘要： Aspects include adding text captioning to a video conference. Multiple conferencing endpoints participate in a video conference. An endpoint locally captures an audio stream and transcribes human speech included in the audio stream into a caption stream. The endpoint can multiplex the caption stream with the audio stream and/or with a captured video stream into a transport stream. The endpoint sends the transport stream to the one or more other conferencing endpoints. To increase reliability and effectiveness, the conferencing endpoint can send the caption stream redundantly. A receiving endpoint can receive and demultiplex the transport stream. The receiving endpoint can coordinate output of the caption stream, the audio stream, and the video stream at corresponding output interfaces.