Multi-Modal Answer Validation for Knowledge-Based VQA

Jialin Wu; Jiasen Lu; Ashish Sabharwal; Roozbeh Mottaghi

Multi-Modal Answer Validation for Knowledge-Based VQA

Jialin Wu, Jiasen Lu, Ashish Sabharwal, Roozbeh Mottaghi

[AAAI-22] Main Track

Keywords
Poster Session 1 @ Blue 4, Poster Session 11 @ Blue 4, Oral Session 11 @ Blue 4, Poster Session 1, Poster Session 11, Oral Session 11

Download Paper

Enter the Virtual Venue

Abstract: The problem of knowledge-based visual question answering involves answering questions that require external knowledge in addition to the content of the image. Such knowledge typically comes in a variety of forms, including visual, textual, and commonsense knowledge. The use of more knowledge sources, however, also increases the chance of retrieving more irrelevant or noisy facts, making it difficult to comprehend the facts and find the answer. To address this challenge, we propose Multi-modal Answer Validation using External knowledge (MAVEx), where the idea is to validate a set of promising answer candidates based on answer-specific knowledge retrieval. This is in contrast to existing approaches that search for the answer in a vast collection of often irrelevant facts. Our approach aims to learn which knowledge source should be trusted for each answer candidate and how to validate the candidate using that source. We consider a multi-modal setting, relying on both textual and visual knowledge resources, including images searched using Google, sentences from Wikipedia articles, and concepts from ConceptNet. Our experiments with OK-VQA, a challenging knowledge-based VQA dataset, demonstrate that \mavex achieves new state-of-the-art results.

Introduction Video

Sessions where this paper appears

Timezone

Poster Session 1

Blue 4

{ "name":"Multi-Modal Answer Validation for Knowledge-Based VQA (Poster Session 1)", "description":"", "startDate":"02-24-2022", "endDate":"02-24-2022", "startTime": "08:45", "endTime": "10:30", "location": "Blue 4", "timeZone": "US/Pacific", "options":[ "Apple", "Google", "iCal", "Microsoft365", "Outlook.com", "Yahoo" ] }

Poster Session 1
Poster Session 11

Blue 4

{ "name":"Multi-Modal Answer Validation for Knowledge-Based VQA (Poster Session 11)", "description":"", "startDate":"02-27-2022", "endDate":"02-27-2022", "startTime": "16:45", "endTime": "18:30", "location": "Blue 4", "timeZone": "US/Pacific", "options":[ "Apple", "Google", "iCal", "Microsoft365", "Outlook.com", "Yahoo" ] }

Poster Session 11
Oral Session 11

Blue 4

{ "name":"Multi-Modal Answer Validation for Knowledge-Based VQA (Oral Session 11)", "description":"", "startDate":"02-27-2022", "endDate":"02-27-2022", "startTime": "18:30", "endTime": "19:45", "location": "Blue 4", "timeZone": "US/Pacific", "options":[ "Apple", "Google", "iCal", "Microsoft365", "Outlook.com", "Yahoo" ] }

Oral Session 11