⚡ New — Kimi K3 is live: bring your own Moonshot key →

Benchmarking Sangam — measured, not marketed.

Consensus routers are easy to hype and hard to compare fairly. Here's how we put Sangam up against OpenRouter Fusion — with a real, money-spending benchmark harness, ablind LLM judge, and a cost model whose assumptions are all on the table. We don't quote a number we haven't measured. Where we only have an estimate, we say so.

Try it yourselfSee the methodology

Honest up front: open-sangam is an open-weight panel — transparent and reproducible, served on BharatRouter credit (your own provider keys are used instead if you've saved them) — and it sits below Fusion on raw frontier quality by design; sangam uses a same-class frontier panel and should land roughly level with Fusion. We don't fight Fusion on the ceiling — we win on the axes it can't follow: openness, reproducibility, cost, and BYOK.

The three engines

Every run puts the same prompts through three consensus engines. Two are ours; one is the router we're measuring against. The cost column is a rate-card estimate(assumptions below) — the harness replaces it with exact billed cost.

EnginePanelResidencyOpennessEst. ₹/req
bharatrouter/open-sangamheroopen-weight · platform-served
GLM-5.3-Flash · Qwen3.6-35B-A3B · Kimi-K2.6 — open-weight, on BharatRouter credit (BYOK optional)
🇮🇳 India available✓ open-weight≈ ₹0.29 / req · ₹0.10 with india_only
bharatrouter/sangamfull frontier
gpt-5 · gemini-2.5-flash · claude-haiku-4.5 → gpt-5-mini
🌐 globalclosed frontierbetween — frontier panel, INR catalog rates
openrouter/fusionfull frontier
Opus · GPT · Gemini + a fuser
🌐 globalclosed frontier≈ ₹1.5–2.0 / req

bharatrouter/sangam is the apples-to-apples Fusion competitor: a same-class frontier panel, but billed at our INR catalog rates — its GPT/Gemini/Claude members need your own keys for those providers. bharatrouter/open-sangam is the open-weight route — transparent, reproducible, and served on BharatRouter credit so it runs for a new org as-is, with India residency available via data_policy — the one to try first.

Where each engine wins

Raw benchmark quality is one axis, not the only one. On a hard reasoning leaderboard, Fusion's frontier panel is hard to top — so we say so plainly. But openness, reproducibility, cost, BYOK, and VPC-deployability are axes a closed frontier router can't follow. Pick the engine your constraints demand.

 open-sangamsangamFusion
India data-residency available (via data_policy)✓ YesNo — frontier panelNo — foreign-deployed providers
Open-weight panel (self-hostable)✓ YesNoNo
BYOK — bring your own keys✓ Yes✓ YesNo
Foreign-deployed frontier model in the loop✓ NoneYes — by designYes — that's its panel
Reproducible (panel composition published)✓ Published✓ PublishedNot disclosed
Runs inside your own VPC (engine-mode)On the roadmap — open weights make it possibleNo — closed panelNo — closed weights
Est. cost / request≈ ₹0.29 · ₹0.10 with india_onlybetween — INR catalog rates≈ ₹1.5–2.0
Raw frontier-quality ceilingLower — open-weight panel, by design≈ Fusion — same-class panelHighest — frontier panel

Read the last two rows together with the rest. If the question is "best possible answer, cost and residency no object," Fusion is a fair pick — and so is our own sangam, at INR rates. If it's "best answer on open weights I can audit, reproduce, and eventually run myself — on our credit or my own keys, with India residency when I need it," that'sopen-sangam.

Cost model

Rate-card estimateNot a measured result. Calculated from published per-token rates under the stated assumptions — the harness replaces every figure here with OpenRouter's exact billed cost.

Under the same workload, open-sangam costs≈₹0.29 per request — about 5–7× cheaper than Fusion. With data_policy: india_only it falls to ≈₹0.10, or16–21× cheaper: the India-resident configuration is also the cheapest one.

  • open-sangam ≈ ₹0.29 / request — five calls (three panelists, a verify pass and a synthesis pass) at BharatRouter's INR catalog rates, debited from your BharatRouter balance. Save your own provider keys and BYOK is used instead, at your provider's rates.
  • …and ≈ ₹0.10 with india_only — the flag dropskimi-k2.6 and keeps the two Krutrim models. kimi-k2.6 alone is64% of the default figure, so residency and cost pull the same way here rather than trading off.
  • Fusion ≈ ₹1.5–2.0 / request — sum of frontier completions at frontier-provider rates, converted at ~₹86/$.
  • bharatrouter/sangam sits between — a frontier panel like Fusion's, but billed at our INR catalog rates rather than raw frontier-provider rates.

Assumptions: ≈400 input + ≈400 output tokens per panel member, with the verify and synthesis passes additionally reading every panel answer; published INR catalog rates (glm-5.3-flash ₹15/₹48, qwen3.6-35b-a3b ₹7/₹28 andkimi-k2.6 ₹91/₹384 per Mtok in/out); Fusion converted at ~₹86/$. Token counts, model mix and FX move the figure — this is a rate-card estimate, not a quote. Recomputed from the catalog on every CI run bysite/test/docs-sangam-cost.mjs, so it cannot silently go stale.

Corrected 2026-09-21. This card previously said "roughly an order of magnitude cheaper" and priced the panel as BYOK-only. The panel has been platform-served since it moved to Krutrim, and at catalog rates the default configuration is 5–7×, not 10×. Only the india_only configuration reaches an order of magnitude — and the page was not describing that one.

When you run the harness, the cost columns stop being estimates: it records the real billed cost OpenRouter reports for each call, so you see what these engines actually charged onyour prompts.

Methodology — the harness

The benchmark isn't a slide; it's runnable code in the repo:bench/sangam-vs-fusion.mjs. It runs the same prompt set through all three engines, has a blind LLM judge pick the best answer, and records latency, tokens, and measured cost. It refuses to run without keys — because it spends real money on real provider calls, there's no fake-data mode.

📋
Shared prompt set

Same prompts, every engine

A fixed set spanning reasoning, India-tax, Hindi, code, and summarization — so each engine answers identical work, not a curated home-field selection.

🙈
Blind best-of-3 judge

The judge can't see who's who

An LLM judge reads the three answers in randomized, anonymized order and picks the best. It never learns which engine produced which answer — no house bias.

🔁
ROUNDS for stability

Repeat, don't trust one shot

Each prompt runs over multiple ROUNDS so a single lucky or unlucky generation doesn't decide the outcome. One judge call is never the verdict — results are aggregated.

📊
Measured, recorded

Latency · tokens · real cost

For every call it logs wall-clock latency, token usage, and the cost the provider actually billed — the numbers that replace this page's estimates.

Run it yourself — supply your own keys (it will spend real money):

BR_KEY=… OPENROUTER_API_KEY=… node bench/sangam-vs-fusion.mjs

One judge model is not the last word — a single judge has its own preferences. Treat the harness as a method for getting comparable numbers under controlled conditions, and read its output as a distribution across prompts and rounds, not a single score.

Measured results — first run

A first pass through the harness — 6 prompts (reasoning, India-tax, Hindi, code, summarization), 1 round, blind best-of-3 judge (gpt-5). A small, directional sample, not a leaderboard — but real numbers, nothing fabricated.

EngineMedian latencyCost / requestBlind quality
bharatrouter/open-sangam~17 s≈ ₹0.29 (rate-card est. — not metered this run)tie 6/6
bharatrouter/sangam~22 s≈ ₹8–13 (est. — frontier panel, not free)tie 6/6
OpenRouter Fusion~55 s≈ ₹13 (measured — $0.89 / 6 calls)tie 6/6

Three honest takeaways: open-sangam was the fastest (~3× quicker than Fusion), ran at ≈₹0.29 a request at catalog rates — its panel is all open-weight and platform-served, far below Fusion's ~₹13 on this run — and the judge scored every prompt a tie — open-weight consensus held its own against a frontier panel on these general questions.sangam is the frontier-panel route, so it is not free: its cost is comparable to Fusion (~₹13), just billed at our INR catalog rates.

These numbers are from 2026-06-17 and the panel has been rebuilt twice since— first to a Groq/Fireworks trio, then on 2026-09-16 to the current Krutrim panel (GLM-5.3-Flash · Qwen3.6-35B-A3B · Kimi-K2.6). The latency and quality figures below describe the original composition, not the one you would call today; treat them as evidence that open-weight consensus can hold its own, not as a measurement of the current engine. A re-run on the current panel is the honest next step. Caveats, kept honest: one round, six prompts, a single judge — ties may partly reflect judge conservatism, and harder reasoning would likely separate the field. Latency is wall-clock with max_tokens 700. On cost: only Fusion (OpenRouter-billed) was directly metered this run; open-sangam's ₹0.29 is the rate-card estimate from the cost model above, and sangam's figure is a rate-card estimate (BR responses don't return a cost field), frontier-class and explicitly not the open-weight rate of open-sangam. The bigger axes — residency, openness, reproducibility — are categorical wins above, not in this table. Re-run it yourself:node bench/sangam-vs-fusion.mjs.

Consensus uplift — same base model

The result above compares whole engines. This one isolates a narrower question: does theconsensus + verify machinery itself add accuracy — separate from just using a bigger model? To find out we pit open-sangam (3-model panel + verifier + synthesizer) against its own lead model, llama-3.1-8b-instruct, runningalone. Same base model on both sides — so any gap is the ensemble machinery, not extra parameters. Grading is objective (every task has a known correct answer, checked by answer-match) — no LLM judge, so no judge bias.8 classic reasoning & arithmetic traps × 5 rounds =40 graded answers. We ran it twice, independently, and publish both — so what you see is a range across real runs, not one lucky shot.

Runopen-sangam (panel + verify + synth)llama-3.1-8b (lead model, alone)UpliftLatency (ens / single)
Run 197.5% (39/40)87.5% (35/40)+10.0 pts~11.8s / ~6.2s
Run 295% (38/40)87.5% (35/40)+7.5 pts~11.8s / ~5.9s

Across two independent runs the ensemble scored 95.0–97.5% against the single model's stable 87.5% — a +7.5 to +10 point uplift on the same base model. That gap is consensus and verification doing their job, not a heavier model. The per-task breakdown (most recent run) shows where: the ensemble rescues traps a single small model slips on — e.g. clock-angle and speed-convert(5/5 ensemble vs 4/5 alone), and the harddigit-7 count (3/5 vs 2/5) where both still struggle — honest about the ceiling.

TaskCategoryCorrect answeropen-sangamsingle
avg-speedreasoning40 km/h (not 45)5/55/5
eggsreasoning245/55/5
widgetsreasoning5 minutes5/55/5
speed-convertquant27.78 m/s5/54/5
days-100logicWednesday5/55/5
digit-7counting203/52/5
clock-anglequant7.5°5/54/5
montyreasoningswitch — 2/35/55/5

How this compares to what the field has published

Consensus isn't ours alone — Together, OpenRouter and Sakana have all published results showing a panel beats its own members. These are different benchmarks on different baselines, so this is not a ranking — read the direction, and note who is independently reproducible.

SystemBenchmarkPanelvs best singleUpliftReproducible?
BharatRouter open-sangam8 reasoning/arithmetic traps · objective grading, no judge95.0–97.5%its own lead model, 87.5%+7.5 to +10 pts✓ harness in repo · 2 runs published
Together MoAAlpacaEval 2.0 (LC win-rate)65.1%GPT-4o, 57.5%+7.6 pts✓ open code · ICLR 2025
OpenRouter FusionDRACO (Perplexity)69.0%Fable 5 solo, 65.3%+3.7 ptsvendor-published
Sakana Fugu UltraSWE-Bench Pro · GPQA-D · LiveCodeBench73.7% · 95.5 · 93.2frontier panel membersleads pool on 10/11✗ vendor-reported, not independently reproduced

Read honestly: Fusion and Fugu lift frontier panels a few points higher (a higher ceiling than our open-weight panel, by design); Together MoA shows open models beating GPT-4o. Ours is the one that isolates the machinery — same base model both sides — and ships a runnable, objective-graded harness with two published runs. Fugu's numbers are vendor-reported and not yet independently reproduced; ours you can reproduce from the repo today. Sources:Together MoA ·OpenRouter Fusion ·Sakana Fugu.

On latency — it's provider-bound, not inherent. The ~11.8s above is the open-weight panel running on platform-served open-weight routes. The consensus pattern is mostly waiting on token generation, so it tracks your provider's speed: in a separate latency probe on Groq's fast inference (open-weight models, 3-model panel + a synthesis pass), the same ensemble shape completed in ~1.25s end-to-end (single ~0.42s). So on production-grade inference, consensus latency stops being the tradeoff — you keep the +10 points without the wait.

Kept honest: the ~1.25s Groq figure is a latency probe on a different panel and prompt than the 40-answer accuracy run above — it measures speed, not quality, and is not conflated with the accuracy numbers. On cost — consensus is not free:a request fans out to roughly 5 calls (a multi-model panel plus verify and synthesis passes), so it spends more tokens than a single call — just on open-weight models. At catalog rates that is ≈₹0.29 a request debited from your BharatRouter balance (≈₹0.10 under india_only), or your own provider's rate if you've saved keys: a real cost, well below a frontier consensus panel like Fusion — but not ₹0. Re-run it yourself: BR_KEY=… node bench/sangam-vs-single.mjs.

Try it yourselfHow Sangam works