Author
Babasaheb Satpute
0 works0 citations
Recent research
- AI & ComputingOpen access
Recent multimodal large language models (MLLMs) have achieved remarkable performance on visual question answering (VQA) and multimodal reasoning tasks. But one overlooked failure case stubbornly persists: when answers require simultaneous synthesis of evidence from natural images...