Submitted to Nature Scientific Reports. Under peer review — not yet published.
Focus
Visual question answering is well studied in high-resource languages and thin for Bangla. This paper studies a hybrid retrieval-augmented (RAG) framework that couples vision with language models under low-resource constraints.
Topics
- Visual question answering (VQA)
- Bengali / Bangla NLP
- Retrieval-augmented generation
- Large language models
Authors and venue are as listed in the Lab record. Claims beyond the submission record wait on peer review and publication.
