A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
-
Updated
Jul 7, 2026 - Python
A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL
Bottom-up attention model for image captioning and VQA, based on Faster R-CNN and Visual Genome
The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".
Oscar and VinVL
[ICCV 2021- Oral] Official PyTorch implementation for Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers, a novel method to visualize any Transformer-based network. Including examples for DETR, VQA.
An efficient PyTorch implementation of the winning entry of the 2017 VQA Challenge.
Visual Question Answering in Pytorch
[CVPR 2021 Best Student Paper Honorable Mention, Oral] Official PyTorch code for ClipBERT, an efficient framework for end-to-end learning on image-text and video-text tasks.
A curated list of Visual Question Answering(VQA)(Image/Video Question Answering),Visual Question Generation ,Visual Dialog ,Visual Commonsense Reasoning and related area.
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
Implementation for the paper "Compositional Attention Networks for Machine Reasoning" (Hudson and Manning, ICLR 2018)
PyTorch implementation for the Neuro-Symbolic Concept Learner (NS-CL).
PyTorch implementation of "Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning"
A lightweight, scalable, and general framework for visual question answering research
To associate your repository with the vqa topic, visit your repo's landing page and select "manage topics."