Hey everyone! I’m building a multimodal RAG pipeline where Mistral OCR annotates images before they go into a vector store with document text.
Issue: Mistral OCR processes images in isolation, so the annotations miss out on critical document context.
Looking for advice on:
Any prompting guides for machine-to-machine image description models to inject context?
Any alternative models or workflows that natively factor in surrounding document context?
Would love to know how you all handle this!
submitted by /u/MediocreAd3005
[link] [comments]