Multimodal RAG – Natural-Language Satellite Image Retrieval

End-to-end development of a multimodal RAG system for natural-language search and automatic captioning of satellite imagery, built on multispectrally adapted vision-transformer models.

Client
Open Beta
Timeframe
2024 – 2025

Key contribution

  • Designed and built the end-to-end multimodal RAG system
  • Implemented the multispectrally adapted, ViT-based retrieval approach
  • Built automated image-captioning generation for Earth-observation contexts

Outcome

Faster, more accessible EO analysis through natural-language search, reducing manual effort in climate monitoring, disaster response and land-use planning.

Focus

  • Multimodal RAG
  • Vision Transformers
  • Earth Observation
  • NLP / LLM
  • Product & Technology Concept