This invention analyzes conversations involving at least two people. It takes in video, sound, and text data for each person's spoken parts. It then processes this data for each person to create individual outputs, which are combined for a specific time window. These combined features are then fed into a machine learning system to generate indicators about the conversation.
Why it matters: While the core technology existed in 2024, the rapid advancements in multimodal foundation models and their accessibility since then could significantly simplify the development and improve the accuracy of the conversation analysis indicators.
AI gives you a few directions you could take this. Pick one, and we check whether your version is different enough to patent, then write the filing.
Reinvent this with AI