Skip to main content
Vincony
AI OSPricingTrust
Log inStart Free
Free Credits
  1. Models
  2. Speed

Speed Benchmarks

Compare inference speed (tokens/sec) across AI providers. Speed-optimized providers like Groq use custom silicon for dramatically faster output.

Fastest: Gemma 2 9B (Groq) at 950 tok/s

Output Speed (tokens/sec)

Swipe to compare
Model Provider Tokens/sec TTFT (ms) CategoryNote
🥇Gemma 2 9B (Groq)Groq95030
Speed-Optimized
Small model + LPU
🥈Llama 3.3 70B (Groq)Groq82045
Speed-Optimized
LPU custom silicon
🥉Llama 4 Maverick (Groq)Groq71055
Speed-Optimized
LPU inference
Mixtral 8x7B (Groq)Groq58060
Speed-Optimized
MoE + LPU
Gemini 2.5 Flash LiteGoogle48050
Standard
Fastest Gemini
Gemini 2.5 FlashGoogle35080
Standard
Optimized for speed
GPT-5 NanoOpenAI26095
Standard
—
GPT-5 MiniOpenAI155180
Standard
—
Grok 3xAI140200
Standard
—
Devstral 2Mistral130170
Open Source
Code-optimized
Claude Sonnet 4Anthropic125230
Standard
—
Gemini 2.5 ProGoogle120250
Standard
—
Gemini 3 Pro PreviewGoogle110280
Standard
—
Llama 4 MaverickMeta100260
Open Source
—
DeepSeek V3DeepSeek95280
Open Source
—
Mistral LargeMistral90240
Open Source
—
GPT-5OpenAI85320
Standard
—
GPT-5.2OpenAI75380
Standard
Heavy reasoning
DeepSeek R1DeepSeek55400
Open Source
Reasoning model
Claude Opus 4Anthropic42450
Standard
Highest quality

Custom Silicon

Groq's LPU (Language Processing Unit) is purpose-built for LLM inference, delivering 5–10× faster output than GPU-based providers.

TTFT Matters

Time-to-first-token (TTFT) affects perceived speed. Speed-optimized providers often respond in under 60ms — instant for users.

Trade-offs

Speed-optimized providers currently support fewer models. Frontier models (GPT-5, Claude Opus) prioritize quality over raw throughput.

These figures are third-party vendor specifications and community-reported numbers — not Vincony's own measured latency. Actual performance varies by prompt length, concurrency, region, and model configuration. We're building a live, measured benchmark to replace them.

Compare Models Side-by-Side Browse All Models
Vincony

Access the world's most powerful AI models through a single, unified platform.

Product

  • All Models
  • Chat
  • Image Generation
  • Video Generation
  • Voice Studio
  • Song Studio
  • All Tools
  • Pricing
  • Integrations
  • API & Developers
  • Download Apps

Solutions

  • Use Cases
  • By Role & Industry
  • Case Studies
  • Testimonials
  • Marketplace
  • Templates
  • Agency Portal
  • White-Label

Resources

  • Help Center
  • Guides
  • Glossary
  • Blog
  • Changelog
  • Feedback
  • Savings Calculator
  • Credits Calculator
  • Plan Recommender

Company

  • About
  • Contact
  • Contact Sales
  • Security
  • Trust Center
  • Bug Bounty
  • System Status
  • Partners
  • Affiliate Program
  • Refer & Earn
  • Brand & Media

Legal

  • Terms of Service
  • Privacy Policy
  • Data Processing Agreement
  • Acceptable Use
  • Cookie Policy
  • Refund & Cancellation
  • Accessibility
  • Sub-processors
  • DMCA & Copyright
Compare AI platforms·Best AI tools·All alternatives·Sitemap

© 2026 VINCONY AI LTD (17047337). All rights reserved.

VINCONY AI LTD · Company No. 17047337 · 3rd Floor, 86-90 Paul Street, London EC2A 4NE, England

GDPR Ready · CCPA Compliant · SOC 2 Aligned · 256-bit Encryption ·

Get weekly AI tips & updates