Llama-3.2-11B-Vision-Instruct

Open multimodal Llama model for image understanding, captioning, and visual QA

BENCHMARKS

Not in the Epoch Capabilities Index. Epoch benchmarks what it can run, so this is a gap in coverage, not a poor result.

RELEASE DATESep 25, 2024
WEIGHTSOpen weights
CONTEXT WINDOW128K
MAX OUTPUT4.1K
REASONINGNo
TOOL CALLINGYes
INPUTStext, image
OUTPUTStext

Specifications from Models.dev. Availability and limits may differ by provider. Benchmark scores come from Epoch AI under CC BY 4.0.