Executive Summary
This report defines a measurement design for multi-model agreement, soft divergence, and material conflict in governed AI decisions. Publication date: 2026-09-10. Last updated: 2026-09-10.
A measured public agreement or divergence rate is not yet available from the corporate evidence corpus. Measurement not yet available for quantitative outcomes; methodology published; benchmark results forthcoming.
The design treats consensus and divergence as decision signals for human review — not as proof of truth. That distinction aligns with NIST AI RMF’s separation of measurement and governance functions[1][2] and with accountability expectations in the OECD AI Principles[3] and ISO/IEC 42001 management-system framing[4].
Related research: taxonomy of multi-model conflict, refusal measurement design, and the program anchor multi-model evaluator methodology. Operational context: consensus vs. divergence, multi-model AI governance, and routing and selection governance.
Key Findings
Findings below are methodological or pending. No measured rates are published in this release.
- A measured public agreement/divergence rate is not yet available from the corporate evidence corpus.
- Agreement among models is a review signal; it is not evidence that an output is accurate or authorized.
- Material conflict (including opposing recommended actions) should escalate to named human authority rather than be averaged.
- NIST AI RMF treats measurement and governance as distinct functions[1].
- Decision Ledger field design specifies the comparison artifacts required for later reconstruction; completeness rates remain pending.
Research Question
Under identical prompts and frozen model versions, how often do eligible models agree, diverge stylistically or factually, or produce decision-critical conflict — and how should those outcomes route into human authorization?
This question is answerable only with a labeled, version-frozen corpus. Until that corpus is released, the answer remains: measurement not yet available.
Dataset and Measurement Status
Dataset size: Not published — measured corpus not yet released.
Methodology document ID: SL-RG-METHOD-2026.1 (methodology only — not a dataset).
Quantitative status: methodology published; benchmark results forthcoming. No sample sizes, percentages, or agreement rates are claimed in this report.
Methodology
Prompt sets are fixed before scoring. Model identity, provider, family, and version are frozen for each run. Outputs are classified into agreement, soft divergence, and material conflict using the taxonomy in the disagreement dataset report.
Materiality criteria distinguish stylistic variance from recommendation-level or policy-level conflict. Classification rules and reproducibility requirements are defined in SL-RG-METHOD-2026.1 and elaborated in the evaluator methodology.
Measured agreement rates, divergence rates, and conflict frequencies: measurement not yet available.
Agreement and Divergence Outcome Classes
The exhibit below is a classification framework for future measurement. It does not report observed frequencies.
| Outcome class | Definition (measurement design) | Typical governance response |
|---|---|---|
| Agreement | Eligible models produce substantively equivalent recommendations under the same prompt and freeze | Still require authorization where outcomes are consequential; do not auto-approve |
| Soft divergence | Differences in style, emphasis, or non-decisive detail without opposing recommended actions | Reviewer notes variance; may select, edit, or request clarification |
| Material conflict | Opposing recommendations, incompatible risk judgments, or policy-incompatible advice | Escalate; preserve all compared outputs; record rationale |
| Incomplete / non-comparable | Refusal, abstention, or failure that prevents a fair multi-model comparison | Route per refusal protocol; see refusal benchmark design |
Source: Smart Logic AI research methodology SL-RG-METHOD-2026.1; measured results not yet published.
Management Implications
Treat multi-model comparison as a control design problem: which conflicts must reach a person, what artifacts must be retained, and who may authorize despite disagreement.
Do not use unpublished rates in board materials, RFPs, or risk registers. Cite this methodology and the pending measurement status until a measured corpus is released.
Connect comparison workflows to Decision Ledger retention guidance in what to record in a Decision Ledger.
Limitations
No public measured corpus is released with this publication. Cross-model rates cannot be inferred from product marketing copy or synthetic demos.
Classification depends on human or adjudicated labels; label drift and prompt sensitivity remain open validity threats until benchmarks are published.
External frameworks cited here set expectations for risk management and accountability; they do not supply Smart Logic AI benchmark scores[1][2][3][4].
How Smart Logic Approaches This
SmartSolo is designed to surface multi-model comparison for human review and to retain path, provenance, and authorization in a Decision Ledger — not to declare consensus as truth.
Corporate research publishes methodology first. Quantitative results will follow only when a measured corpus meets the reproducibility bar in SL-RG-METHOD-2026.1.
For program design, see multi-model AI governance and multi-model AI orchestration.
References
Authoritative sources cited for standards and methodology claims. Inline markers link here.
Citation kit
Report title: When AI Models Agree—and When They Don't: A Multi-Model Decision Benchmark
Publisher: Smart Logic AI
Publication date: 2026-09-10
Last updated: 2026-09-10
Canonical URL: https://www.smartlogicusa.com/research/multi-model-agreement-divergence-benchmark
Methodology URL: https://www.smartlogicusa.com/research/multi-model-evaluator-methodology
Dataset / version identifier: Not assigned — measured corpus not yet published
Method document ID: SL-RG-METHOD-2026.1
Data coverage: Measured public corpus not yet released; methodology coverage begins 2026-09-10.
Suggested citation: Smart Logic AI. “When AI Models Agree—and When They Don't: A Multi-Model Decision Benchmark.” 2026. https://www.smartlogicusa.com/research/multi-model-agreement-divergence-benchmark.
See governed AI execution in a live workflow
Review how SmartSolo coordinates multiple AI models, routes human authorization, and preserves the decision record.