Consensus & Differences Engine

Consensus and Difference Engines explained.

A clear look at how consens.io combines model answers, checks disagreements, and checks the result in a separate judge pass.

One question · 6 models What is the longest river in the world?

68 Strong agreement 1 minor detail · disputed: where the Amazon starts, 6 models compared

Consensus Answer

The Nile is the conventional longest river, but the Amazon has a credible claim. By convention the Nile leads at about 6,650 km Whether the Amazon overtakes it depends on where its mouth is drawn, and a disputed 2007 survey draws it further out than most references do.

Highlights show model agreement and disagreement, not a fact-check. Open the original model answers for context.

See all differences
Contradiction · minor detail Is the Nile definitively the longest river?
ChatGPT, Mistral, DeepSeek
The Nile is the settled answer at about 6,650 km.
Claude, Gemini, Grok
The 2007 Amazon survey is a credible challenger.

How it works

The engine runs in four readable steps.

The answer models respond independently, your selected Consensus Engine writes the synthesis, and two judges analyze differences and sentence coverage. In Consensus mode, an optional source check then examines eligible factual contradictions in Consensus Chat.

POST /consensus Question, included model answers, excluded models, sources, and selected engine

The backend authenticates the user, checks tier and usage rules, caps client-sent text, and requires at least two included model answers.

1

Consensus Engine

stream_consensus()

  1. Build synthesis prompt Drop empty or excluded answers, shuffle the remaining responses, anonymize them as Expert A/B/C, and attach compact source provenance.
  2. Resolve selected engine The dropdown value maps to a provider model, including standard, direct model IDs, or Pro aliases.
  3. Stream the synthesis The selected engine writes the user-facing answer. Provider failures are retried twice, then a fallback family is tried when a key is available.
  4. Gate the second pass If the final consensus is empty or an error text, the comparison is skipped so the judge never analyzes an error message.
2

Differences Engine

stream_differences()

  1. Build judge prompt Use the consensus answer plus included source answers, cap each answer for the judge pass, shuffle, and anonymize as Model A/B/C.
  2. Pick the judge The judge tier follows the selected Consensus Engine; the family follows a fixed priority (OpenAI first, then Gemini), so it can be the engine's own.
  3. Request strict JSON The Differences Judge returns difference cards with positions, severity, verification hints, and the closest source model.
  4. Cover every sentence in parallel A second, cheaper judge runs at the same time on the binding list of numbered sentences and must answer for each one: support, contradiction, or not addressed.
  5. Verify and score The server repairs JSON if safe, maps labels back, verifies quotes and anchors, computes the agreement score, and records judge metadata.
SSE final payload consensus_response + differences_data + usage/share metadata

The UI renders one readable consensus from this payload, marks the anchored passages inline, and keeps the agreement score, the contradiction entries and the judge note attached to it.

01

Up to six model families answer independently.

Choose from OpenAI, Claude, Gemini, Mistral, DeepSeek, Grok, Kimi, GLM, and Meta Muse. The selected families receive the same prompt, and their original answers stay separate, so the synthesis never replaces the raw evidence.

02

The selected Consensus Engine writes the synthesis.

The engine you choose in the dropdown reads only the included answers, ignores excluded ones, and writes one response that preserves common ground plus necessary caveats.

03

Two separate judges check the consensus.

The Differences Judge works out the substantive disagreements. In parallel, the Coverage Judge goes through the numbered sentences and records, for each one, which models support it, contradict it, or never addressed it. Both run as separate calls after the synthesis, and the judge model can come from the same family as the engine.

04

The UI marks the evidence inside the answer.

The final payload contains the consensus text, agreement score, claim badges, contradiction entries with their anchors, and judge metadata, so each disputed passage can be marked exactly where it stands.

Model roles

Different models do different jobs.

The process is intentionally split. The model that writes the consensus is not treated as the only authority on whether the consensus is reliable.

Answer layer

OpenAI, Claude, Gemini, Mistral, DeepSeek, Grok, Kimi, GLM, and Muse

These are the source answers. They are queried in parallel, shown to you directly, and then passed into the consensus step if they are not excluded.

Synthesis layer

The engine chosen in the Consensus picker

This model reads the included source answers and writes the combined response. Pick a preset (Daily, Balanced, High Quality) or choose a specific model via Custom. Standard and Pro engines can both be used.

Analysis layer

Two analysis judges, plus an optional source check

The Differences Judge works out the contradictions; the Coverage Judge belongs to the same pass but has one narrow job, and therefore runs on a cheaper model: a verdict for every sentence, with none allowed to be skipped. Both follow the same judge priority, OpenAI first, then Gemini.

In Consensus mode with Check contradictions on, an additional judge checks major factual disagreements against sources already supplied by both model positions. Results appear beside each contradiction, with original passages, source links and relevant dates and conditions. Background checks have explicit budgets; unavailable and omitted checks are identified. This is not a complete fact-check of the consensus. The answer, model positions and agreement score stay unchanged.

Differences Engine

The second pass turns disagreement into structured evidence.

After the consensus text is written, consens.io runs a separate analysis pass. Its job is not to rewrite the answer, but to explain how the source models support, qualify, or contradict it.

01

Inputs are filtered and anonymized.

Only non-empty, non-excluded model answers are used. Each answer is capped before analysis and then renamed to neutral labels like Model A, Model B, and Model C.

Why: anonymization and shuffled order reduce provider-name and position bias in the judge pass.

02

Both judges must return strict JSON.

The Differences Judge is asked for difference cards, contradiction severity, verification hints, and the model closest to the consensus. The Coverage Judge is asked for one entry per sentence id, with a stance for every model - and the server rejects a reply that leaves an id out, asking for the missing ones again.

Types: contradiction for incompatible facts or conclusions, emphasis for different focus or missing context.

03

How the judge is chosen.

The judge tier follows the selected Consensus Engine: standard consensus uses a standard judge, Pro consensus tries a Pro judge. The family follows a fixed priority, OpenAI first, then Gemini, even when the engine belongs to the same family.

Fallback: primary judge, retry, then the next available family; Pro attempts can fail open to a standard judge.

04

The server verifies and scores the result.

JSON is repaired if safely possible, model labels are mapped back, quotes are checked against the original answers, and unverifiable quotes are removed.

Output: agreement score, level, major/minor contradiction counts, emphasis counts, and judge metadata.

Structured payload

What comes out of the Differences Engine

claimsOne entry per checkable sentence: the consensus anchor, the models that support it, the models that contradict it with their quotes, and how far it is covered at all.
differencesContradiction or emphasis entries with positions and verification hints, each with a consensus_anchor: the verbatim passage in the consensus answer that the disagreement is about.
agreementA 0-100 score computed from claim support and weighted disagreement penalties.
judgesThe provider, model, and tier that actually delivered the analysis.
64 Partial agreement 1 critical · disputed: the condition attached to it, 6 models compared Analysis by Gemini (Pro) separate judge pass

Differences

Contradiction · critical Does the caveat change the recommendation?
ChatGPT, Mistral, DeepSeek
The recommendation holds as stated.
Claude, Gemini, Grok
It only holds under a specific condition.
Worth verifying: check whether the condition applies to your case before acting on the answer.

What you see

A useful answer without pretending everything is settled.

There is one answer to read. Where the judge found a verified disagreement, the passage that carries it is marked in place, and the full analysis stays one click below.

82 Strong agreement 1 minor detail · disputed: how strong the evidence is, 6 models compared Analysis by Mistral separate judge pass

Consensus Answer

Most models agree on the main recommendation.

The answer combines the repeated points and removes duplicated wording How strong the underlying evidence is turns out to be the open question, because the sources rate the measurement quality differently.

Every checkable sentence was compared against each model: green means supported, amber means the models are split, red marks a contradiction, and grey means they differed only in detail — or that too few of them addressed the sentence at all. Models agreeing is not proof.

See all differences
Contradiction · minor detail How strong is the evidence?
ChatGPT, Gemini, DeepSeek
The evidence is settled and supports the recommendation.
Claude, Mistral, Grok
A caveat about measurement or source quality applies.
Analysis by a judge model, run as a separate pass after the consensus.

Why this helps

Consensus is a reading aid, not a magic truth label.

Less blind trust

One confident answer is not enough

Seeing several model families makes weak spots easier to notice before you act on an answer.

More signal

Agreement becomes visible

Repeated claims are separated from one-off claims, so you can scan the strongest common ground first.

Clear caveats

Disagreement stays in view

Contradictions are not averaged away. They are shown as specific points to inspect or verify.

Optional · Consensus Chat

What a contradiction source check tells you.

A factual dispute is checked against sources already attached to the model positions. Open the result for original passages, links and the conditions under which each claim holds.

Evidence for a position

The sources may support one position, show that the models refer to different dates or conditions, or conflict with one another. The evidence is shown with the result; original passages are not translated.

No checkable contradiction

The analysis found no suitable factual dispute to check. Preferences and recommendations are excluded. This status does not mean the answer was verified or that every model agrees.

The check is incomplete

Running, switched off, failed analysis, inaccessible sources, insufficient evidence and budget omissions have separate statuses. None is presented as a successful verification.

Source checks leave the answer, model positions and agreement score unchanged. Highlights can be hidden for easier reading; the findings stay available. Resolve is a separate action.

FAQ

Common questions.

Can I turn the source check off?

Yes. Check contradictions starts on in Consensus mode. Compare has no consensus to check, so the control is hidden there; your saved choice returns when you choose Consensus again. Switch it off under the input, in the (+) menu, or in Settings → Runs. Your choice is saved for the next chat run. Watches, Topics and API runs do not perform this optional check. The check uses existing sources from both sides of major factual disagreements. Results distinguish support for a position, different conditions, conflicting sources and insufficient evidence. Preferences and recommendations are excluded. Inaccessible documents and budget omissions get explicit statuses. This does not verify the entire answer. Consensus, claims and differences still run when it is off.

Is the consensus just the majority answer?

No. The engine writes a synthesis from the included model responses. Agreement matters, but the output is written as a readable answer with caveats, not as a simple vote.

Can I choose the Consensus Engine?

Yes. In the app you pick a preset (Daily, Balanced, or High Quality) or select a specific consensus model via Custom. If the resulting engine is a Pro engine, the differences analysis tries to use a Pro judge.

Is the judge from a different model family?

Not necessarily. The judge follows a fixed priority (OpenAI first, then Gemini), so it can share a family with the Consensus Engine or with one of the answers. It sees the answers anonymized and shuffled, its quotes are checked verbatim against them, and the agreement score is counted from its sentence-by-sentence verdicts rather than given as a grade.

What happens if the Differences JSON is malformed?

The backend extracts or repairs a complete JSON object when it can do so safely. If the output still cannot be parsed, raw JSON is not shown to the user.

Try the flow

Ask once, compare the answers, then read the consensus with its caveats.

Open app