Llama 3.2 90B Vision Instruct
Llama 3.2 90B Vision Instruct is Meta's flagship multimodal model from the Llama 3.2 generation, combining a large-scale language backbone with vision encoding for high-quality image and text understanding. As the largest model in the Llama 3.2 family, it brings substantially more reasoning depth to visual tasks than smaller vision-language alternatives.
It is instruction-tuned and safety-aligned, suited for complex multimodal workflows that demand accurate image interpretation combined with nuanced language generation — such as detailed document analysis, scientific figure interpretation, or sophisticated visual reasoning chains. Its open-weight release makes it available for large-scale self-hosted enterprise deployments and research.
Key Features
90B parameters providing strong multimodal reasoning depth
High-fidelity image understanding combined with advanced language generation
Instruction-tuned for complex visual Q&A and multi-step reasoning
Open-weight release supporting self-hosted and fine-tuned deployments
Handles complex documents, diagrams, charts, and natural scene images
Safety-aligned for deployment in user-facing applications
Ideal Use Cases
Detailed scientific or technical document and figure analysis
High-accuracy OCR and structured data extraction from scanned forms
Complex visual reasoning tasks in research and diagnostic workflows
Enterprise knowledge management tools processing mixed image-text content
Multimodal content moderation requiring nuanced visual understanding
Example Prompts for Llama 3.2 90B Vision Instruct
Technical Specifications
| Provider | Meta |
| Category | Text |
| Modality | Text -> Text |
Frequently Asked Questions
Try Llama 3.2 90B Vision Instruct now
Start using Llama 3.2 90B Vision Instruct instantly — 100 free credits, no credit card required. Access 750+ AI models through one platform.