Llama-3.2-11B-Vision-Instruct
Open multimodal Llama model for image understanding, captioning, and visual QA
BENCHMARKS
Not in the Epoch Capabilities Index. Epoch benchmarks what it can run, so this is a gap in coverage, not a poor result.
RELEASE DATESep 25, 2024
WEIGHTSOpen weights
CONTEXT WINDOW128K
MAX OUTPUT4.1K
REASONINGNo
TOOL CALLINGYes
INPUTStext, image
OUTPUTStext
Specifications from Models.dev. Availability and limits may differ by provider. Benchmark scores come from Epoch AI under CC BY 4.0.