Multimodal RAG – Natural-Language Satellite Image Retrieval
End-to-end development of a multimodal RAG system for natural-language search and automatic captioning of satellite imagery, built on multispectrally adapted vision-transformer models.
Key contribution
- Designed and built the end-to-end multimodal RAG system
- Implemented the multispectrally adapted, ViT-based retrieval approach
- Built automated image-captioning generation for Earth-observation contexts
Outcome
Faster, more accessible EO analysis through natural-language search, reducing manual effort in climate monitoring, disaster response and land-use planning.
Focus
- Multimodal RAG
- Vision Transformers
- Earth Observation
- NLP / LLM
- Product & Technology Concept