Llama Guard 3 8B
Llama Guard 3 8B is a safety-focused classifier from Meta built on top of the Llama 3 architecture. Rather than generating free-form text, it is purpose-trained to evaluate whether a given conversation turn — an input prompt or a model-generated response — falls into predefined categories of harmful or policy-violating content.
Meta designed Llama Guard 3 to align with the MLCommons taxonomy of hazard categories, making it suitable as a moderation layer in both consumer-facing and enterprise LLM deployments. At 8 billion parameters it offers meaningful classification accuracy while remaining deployable on standard inference hardware alongside a primary generation model.
Key Features
Classifies both user inputs and model outputs for policy violations
Covers MLCommons hazard taxonomy categories (violence, hate, CSAM, etc.)
Returns structured labels enabling downstream routing or filtering
Open weights for integration directly into existing inference stacks
Supports multi-turn conversation context for more accurate moderation
Fine-tunable on custom safety taxonomies for specific deployment needs
Ideal Use Cases
Real-time content moderation layer for LLM-based chat products
Automated red-teaming pipelines to evaluate model safety posture
Policy compliance screening before displaying AI-generated content
Building responsible-AI guardrails in enterprise AI platforms
Audit logging and flagging of unsafe conversations for human review
Example Prompts for Llama Guard 3 8B
Technical Specifications
| Provider | Meta |
| Category | Text |
| Modality | Text -> Text |
| Context Window | 128K tokens |
Frequently Asked Questions
Try Llama Guard 3 8B now
Start using Llama Guard 3 8B instantly — 100 free credits, no credit card required. Access 750+ AI models through one platform.