This invention describes a system that translates live audio during a video or audio call over a network. When a user on one device receives speech in a foreign language from another person, the system first figures out what language the user is currently speaking or listening in for that specific conversation. If this conversation language is different from the device's usual language, the system translates the incoming speech into the conversation language and presents it to the user. The system can determine the user's current conversation language by analyzing their facial expressions, location, or the sounds they are making locally.
Why it matters: Filed before the significant advancements in real-time neural machine translation and on-device machine learning for audio and visual processing. The capabilities for accurate, low-latency language identification from diverse inputs and high-quality translation have matured considerably since 2019.
AI gives you a few directions you could take this. Pick one, and we check whether your version is different enough to patent, then write the filing.
Reinvent this with AI