This invention describes a method for an electronic device to understand what a user is talking about in an image. It uses one AI model to analyze the objects in a displayed image and another AI model to understand a user's spoken words. By comparing the AI's understanding of the image and the voice, the system identifies a specific object the user is referring to and provides descriptive tags or keywords for it, which are then shown alongside the image.
Why it matters: Filed before large multimodal models (LMMs) became widely available. The integration and comparison of separate image and voice AI models described in 2019 can now be streamlined and enhanced by modern LMMs, which natively handle both inputs.
AI gives you a few directions you could take this. Pick one, and we check whether your version is different enough to patent, then write the filing.
Reinvent this with AI