Skip to main content
Vincony
AI OSPricingTrust
Log inStart Free
Free Credits
  1. Resources
  2. Compare Tool Guide
Back to blog
Tutorials

The Compare Hub Guide: Blind Vote, Bracket, Disagreement Map, and Leaderboard

Vincony TeamJune 6, 20267 min read

# The Compare Hub Guide: Blind Vote, Bracket, Disagreement Map, and Leaderboard

Vincony's Compare tool does a lot more than place two chat windows side by side. It's a full structured-comparison workbench — one destination that absorbed what used to be six separate tools (Taste Test, Tournament, Disagreement Map, Model Race / Trust Arena, Model Rankings, and Compare Lab). The result is a single tabbed hub with five panels, each designed for a different kind of question.

This guide focuses on the four structured-comparison panels: Blind Vote, Bracket, Disagreement, and Leaderboard. If you're looking for the freeform Side-by-side chat panel (where you type to two models simultaneously and read both responses), that workflow is covered separately. Here we're talking about the modes that enforce rules — hidden identities, elimination rounds, consensus scoring — so your judgment stays honest.

---

How the hub is organised

Open Compare and you'll see five tabs across the top of the page:

  • Side-by-side — freeform parallel chat (not this guide's focus)
  • Blind vote — three anonymous models, one prompt, one vote
  • Bracket — eight models, elimination tournament, one champion
  • Disagreement — five models answer a factual question; a meta-model scores where they agree and where they diverge
  • Leaderboard — blind A/B arena with ELO ratings and a Speed Race sub-mode

You can also jump straight to any panel from a URL: add ?panel=blind-vote, ?panel=bracket, ?panel=disagreement, or ?panel=leaderboard to the Compare route.

All four structured panels require an account. Free plan users can access Blind Vote and Disagreement; Bracket and the full Leaderboard arena are available on Starter and above. The Power plan unlocks the Trust Arena sub-mode inside Leaderboard.

---

Panel 1 — Blind Vote

Best for: Discovering which model you actually prefer, stripped of brand loyalty.

The Blind Vote panel runs three randomly selected AI models against the same prompt and shows you the responses labelled only A, B, and C. You vote before identities are revealed. This eliminates the "I already trust GPT-5" bias and forces you to judge on writing quality alone.

Step-by-step:

  • Go to Compare and click the Blind vote tab.
  • Type your prompt in the text area. Six sample prompts are shown as chips — click any to load it instantly. Prompts that work well are open-ended enough to show style differences: explanations, persuasive writing, code tasks, or opinion questions.
  • Check the credit cost shown next to the Run Taste Test button. The panel pulls three models simultaneously, so the cost is the sum of all three.
  • Click Run Taste Test. All three cards stream their responses in parallel. A live elapsed-second timer ticks on each card while it generates.
  • Once all three finish, a prompt bar appears: "Vote for the best response to reveal model names." Click Vote for A, Vote for B, or Vote for C on the card you prefer. You can also click anywhere on the card.
  • After voting, a Reveal Models button appears. Click it to see exactly which model wrote each response, plus word count and generation time per card.
  • Post-reveal options: New Taste Test (fresh prompt and fresh models), Same Prompt, New Models (re-shuffle the lineup), or Same Prompt, Same Models (re-run to check consistency).

Your completed votes accumulate in a personal Your Leaderboard section below the main panel — expand it with the toggle to see which models have won your votes over time.

Tips: - Use Shuffle Models in the header before running to get a completely different trio. - The panel deliberately picks models from different providers when possible, so you're comparing across families, not just versions of the same model. - Try the same prompt three or four times with different model sets. A model that consistently wins on your use case is worth noting on the models page.

---

Panel 2 — Bracket

Best for: Settling once and for all which model wins on a specific task type.

The Bracket panel runs a March-Madness-style elimination tournament. Eight named models — GPT-5.2, Claude Sonnet 4.5, Gemini 3 Flash, DeepSeek R1, GPT-5 Mini, Mistral Large, Claude Sonnet 4, and Grok 4 — compete head-to-head across three rounds (Quarter-Finals, Semi-Finals, Final). Every match is a blind two-response vote, and the model you pick advances.

Step-by-step:

  • Click the Bracket tab on Compare.
  • The eight competing models are listed as badges below the prompt field so you know who's in the draw before you start.
  • Type your prompt. The same prompt runs for every match in the tournament, so choose something representative of the work you actually do. Click Start Tournament (costs approximately 16 credits for all matches).
  • The first quarter-final match loads. Two anonymous response cards appear labelled Response A and Response B, with models hidden. Read both, then click the card you prefer.
  • The winner advances, and the next match loads automatically. Work through all four quarter-finals, then the two semi-finals, then the final.
  • A champion screen appears with a trophy and share button once all rounds complete. You can copy or share the result ("GPT-5.2 won my AI Tournament on Vincony!") directly from the page.
  • Click New Tournament to run a fresh draw with a new prompt.

A compact bracket visualisation below the active match tracks which round you're in and shows checkmarks on completed matchups, so you always know where you are in the draw.

Tips: - The bracket randomises seedings on each start, so the same two models won't always meet in the final. - Use a task that has a right answer, like a specific coding problem or a translation, to get the most signal. Style-heavy prompts (poetry, marketing copy) work too but tend to be more subjective. - Because every match is blind, your vote reflects the response quality, not the model's reputation — which is the entire point.

---

Panel 3 — Disagreement

Best for: Factual research, contested topics, and calibrating how confident you should be in any single model's answer.

The Disagreement panel sends your question to five models simultaneously — GPT-5.2, Claude Sonnet, Gemini 3 Pro, DeepSeek V3.2, and Mistral Large — then runs a sixth model as a meta-analyst to score consensus. The output is structured around three tiers: full consensus (all models agree), partial agreement (most models agree), and outright disagreement (models diverge). An overall confidence percentage is included.

Step-by-step:

  • Click the Disagreement tab on Compare.
  • Enter a factual or analytical question. The panel is built for questions where the answer matters — not creative tasks. Good examples: "What causes inflation?", "Is intermittent fasting effective for weight loss?", "What are the main arguments for and against nuclear energy?"
  • Press Enter or click Map Disagreements (costs approximately 8 credits). All five models query in parallel.
  • Five response cards appear in a grid as results stream in, each colour-coded by model. Cards show the response text, truncated for readability.
  • Once all five finish, the meta-analysis panel slides in below: a structured breakdown showing what the models collectively agree on, where the majority aligns, and where specific models diverge. The confidence score gives you a single-number signal for how settled the question is.
  • Use the Share button in the analysis panel to copy or native-share a compact summary of the result.

Tips: - This panel shines on questions where you've already received one confident answer from a single model and want a second opinion at scale. If the disagreement panel returns 95% consensus, that answer is probably reliable. If it returns 40%, treat the topic as contested and do additional research. - Avoid prompts that are purely creative or highly subjective — the meta-analysis works best when there's a factual ground truth to converge on or diverge from. - The analysis always names which model deviated, so you can note if a specific model consistently hedges or takes a contrarian position on your topic area.

---

Panel 4 — Leaderboard

Stay ahead in AI

Get our weekly AI insights — tips, model comparisons, and guides delivered to your inbox.

No spam, unsubscribe anytime.

Best for: Building an honest, data-driven ranking of models based on your own votes and community votes — with ELO ratings that move after every match.

The Leaderboard panel has two sub-modes accessible via an inner tab row: Trust Arena and Speed Race.

Trust Arena (default)

Trust Arena is a structured blind A/B test that feeds results into a live ELO rating system. Every vote you cast moves the ELO scores of both models up or down. Over time, the leaderboard reflects genuine preference across thousands of real queries.

Step-by-step:

  • Click the Leaderboard tab on Compare. You'll land on the Trust Arena sub-tab.
  • Pick a category from the category picker (coding, writing, reasoning, general, etc.) to keep your votes grouped meaningfully.
  • Toggle Public on if you want your prompt and results to appear in the Community Prompts feed below, visible to other users who might want to try the same comparison.
  • The panel shows two randomly selected anonymous models. Click New Matchup to re-roll the pair if you want a different combination.
  • Type your prompt in the text area (Ctrl/Cmd+Enter to start). Click Start Arena. Both models stream responses in parallel, labelled only Model A and Model B, with elapsed time and word count shown live.
  • Once both finish, the voting bar appears with three options: Model A, Tie, or Model B. Vote for the response you found more useful.
  • After voting, both cards reveal the actual model names. A reveal card shows the ELO change for each model — the winner gains rating points, the loser loses them, weighted by the gap in their pre-match ratings.
  • Click Next Match to queue a new matchup immediately, or scroll down to Community Prompts to browse public results and load a prompt someone else ran.

Speed Race sub-mode

Speed Race is less about quality judgment and more about benchmarking raw throughput. You select two to five models (default: Gemini 3 Flash, GPT-5, Gemini 3 Pro — up to five), run them on the same prompt, and watch them stream simultaneously. The first model to finish gets a crown badge. All responses are ranked 1st through 3rd by completion order with gold/silver/bronze badges.

This mode is useful for latency-sensitive workflows where you need to know which model will respond fastest on a given task type.

Tips for Leaderboard: - Trust Arena votes compound — run it regularly with prompts from your actual work to build a personal leaderboard that's tuned to your use cases. - Use the category picker consistently. Mixing coding and creative prompts in the "general" bucket produces noisy rankings. - The community feed is a shortcut: if you're not sure what to test, browse public prompts that other users ran in your category and load one directly. - Power plan users get access to Trust Arena; Speed Race is available on Pro and above. Check pricing if you hit the gate.

---

Which panel should you use?

GoalPanel
Unbiased preference test across 3 modelsBlind Vote
Crown a champion across 8 models on a specific taskBracket
Check how reliable a factual answer actually isDisagreement
Build a long-term ranked leaderboard with ELOLeaderboard (Trust Arena)
See which model responds fastestLeaderboard (Speed Race)

---

A note on credits

Get the step-by-step checklist

Free — delivered to your inbox instantly.

Structured comparisons use more credits than a single chat because they're running multiple models simultaneously. The UI always shows the credit cost before you confirm, so there are no surprises. If you're on a Free plan and running out of credits mid-session, consider Starter or Pro for higher monthly allowances — details at pricing.

---

Try it now

All four panels live at Compare. Start with Blind Vote if you've never used it — run your single most common prompt type, vote honestly, and see which model you've been sleeping on. Then graduate to Bracket when you want a definitive head-to-head result, and Leaderboard when you're ready to contribute to (and benefit from) a community-built ranking that gets more accurate with every vote.

The more you use it, the better your personal data gets. Your votes are yours — and they're already telling you something your instincts might not.

Actions

Related Articles

Tutorials

How to Build a Custom AI Chatbot for Your Website in Under an Hour

Feb 13, 2026

Tutorials

AI Image Generation Tips & Tricks: From Prompt to Pixel-Perfect Output

Feb 20, 2026

Tutorials

AI Video Generation Tutorial: Create Professional Videos from Text Prompts

Feb 21, 2026

On this page
Vincony

Access the world's most powerful AI models through a single, unified platform.

Product

  • All Models
  • Chat
  • Image Generation
  • Video Generation
  • Voice Studio
  • Song Studio
  • All Tools
  • Pricing
  • Integrations
  • API & Developers
  • Download Apps

Solutions

  • Use Cases
  • By Role & Industry
  • Case Studies
  • Testimonials
  • Marketplace
  • Templates
  • Agency Portal
  • White-Label

Resources

  • Help Center
  • Guides
  • Glossary
  • Blog
  • Changelog
  • Feedback
  • Savings Calculator
  • Credits Calculator
  • Plan Recommender

Company

  • About
  • Contact
  • Contact Sales
  • Security
  • Trust Center
  • Bug Bounty
  • System Status
  • Partners
  • Affiliate Program
  • Refer & Earn
  • Brand & Media

Legal

  • Terms of Service
  • Privacy Policy
  • Data Processing Agreement
  • Acceptable Use
  • Cookie Policy
  • Refund & Cancellation
  • Accessibility
  • Sub-processors
  • DMCA & Copyright
Compare AI platforms·Best AI tools·All alternatives·Sitemap

© 2026 VINCONY AI LTD (17047337). All rights reserved.

VINCONY AI LTD · Company No. 17047337 · 3rd Floor, 86-90 Paul Street, London EC2A 4NE, England

GDPR Ready · CCPA Compliant · SOC 2 Aligned · 256-bit Encryption ·

Get weekly AI tips & updates