Sonar Small
Sonar Small is Perplexity's compact web-grounded model built on a Llama 3.1 base, designed for fast, cost-efficient question answering with real-time internet retrieval. It queries the live web at inference time and synthesizes results into concise, cited responses, making it practical for applications that need current information without the compute cost of larger variants.
Its smaller footprint makes it well-suited for high-throughput pipelines, mobile-adjacent applications, or any scenario where latency and cost must be minimized while still maintaining access to fresh web content. It handles factual lookups, quick summaries, and news queries effectively.
Key Features
Real-time web search grounding at inference time
Inline source citations for answer transparency
Lower latency and cost compared to larger Sonar variants
Llama 3.1 base with Perplexity's retrieval stack
Handles factual queries, news, and current-events lookups
Suitable for high-throughput API usage
Ideal Use Cases
High-volume factual Q&A with live web sources
Cost-sensitive chatbot backends needing current information
News headline summarization at scale
Quick product or pricing lookups requiring fresh data
Developer prototyping of search-grounded features
Example Prompts for Sonar Small
Technical Specifications
| Provider | Perplexity |
| Category | Search |
| Modality | Text -> Text (web-grounded) |
| Context Window | 127,072 tokens |
Frequently Asked Questions
Try Sonar Small now
Start using Sonar Small instantly — 100 free credits, no credit card required. Access 750+ AI models through one platform.