This invention describes a system that lets you select specific objects shown in a video by using sounds from that video. It works by first analyzing the video's audio to create unique sound fingerprints using a hashing function, and simultaneously identifying objects in the video using a trained AI. When a user provides a sound clip, the system matches it to a fingerprint, finds the corresponding video object, and makes it interactive.
Why it matters: Since 2023, the rapid evolution of computer vision and audio processing models, particularly pre-trained neural networks, has significantly improved the accuracy and reduced the development effort for identifying objects and generating audio identifiers. This makes the core components of this system more robust and accessible now.
AI gives you a few directions you could take this. Pick one, and we check whether your version is different enough to patent, then write the filing.
Reinvent this with AI