Multimodal Model
A multimodal model works with more than one kind of information, such as images and text. Which combinations it accepts or produces depend on the particular model.
[Liu et al.]In practice · hypothetical example
A museum assistant receives a photo and a written question about the pictured object.
[Liu et al.]A little deeper
LLaVA is one example: it connects a vision encoder to a language model. Supporting images and text does not imply support for every other modality. [Liu et al.]
A common mix-up
Multimodal means every kind of media is supported.
A model’s supported inputs and outputs must be checked individually. [Liu et al.]
Does image-and-text support guarantee audio input?
Sources & editorial notes
Evidence: supported. Primary-source support for this scoped entry; publication approved by the project owner.
- Visual Instruction Tuning ↗ (opens in new tab)Liu et al. · Publication date unknown
Relevant section: Abstract
Last editorial review: 2026-09-13 by project-owner.
First observed in this corpus: Unknown.
Revision history
Revision 2 · Created 2026-09-13 · Updated 2026-09-13
Project owner approved the current content for publication. Existing evidence scope and limitations remain applicable.