This invention describes a way for an intelligent assistant to answer questions by pulling out both text and relevant images from digital documents. It first finds a textual answer from the most important parts of the document, then extracts images from those same pages. The system then identifies the best images to include by comparing them to the text answer, often using captions, titles, or text found within the images themselves, especially from software documentation with screenshots.
Why it matters: The invention was filed just as multimodal AI models were emerging. The claims' focus on matching textual answers to images, especially screenshots, using semantic similarity and OCR is now significantly more powerful and accurate with the advanced multimodal LLMs available in 2026.
AI gives you a few directions you could take this. Pick one, and we check whether your version is different enough to patent, then write the filing.
Reinvent this with AI