Executive Summary
This document is the methodological anchor for Smart Logic AI’s multi-model governance research program (SL-RG-METHOD-2026.1). Publication date: 2026-09-10. Last updated: 2026-09-10.
It defines prompt-set construction, model freezing, disagreement classification, materiality criteria, completeness rules, and reproducibility requirements for dependent reports. It is not a scored dataset and contains no fabricated benchmark scores.
Dependent studies: agreement/divergence, disagreement taxonomy, traceability, and override study. Authority context: multi-model governance, consensus vs. divergence, model vs. decision governance, and framework vs. platform.
Key Findings
Methodological conclusions only — no scores.
- Quantitative benchmark results across the program remain forthcoming; this document publishes method, not measured rates.
- Reproducible multi-model evaluation requires frozen prompts, frozen model versions, and explicit disagreement labels before any rate is publishable.
- Materiality criteria separate soft divergence from decision-critical conflict so escalation policy can be tested later without inventing frequencies.
- NIST AI RMF separates Govern, Map, Measure, and Manage functions — measurement without governance context is incomplete[1][2].
- OECD AI Principles and ISO/IEC 42001 motivate accountable, documented AI management practice; OMB M-25-21 reinforces federal governance expectations for agency AI use[3][4][5].
Methodology Overview
Program reports share a common evaluation spine: (1) define the decision context and risk class; (2) freeze the prompt inventory; (3) freeze eligible models and versions; (4) capture raw outputs and human actions; (5) label disagreement and refusal classes; (6) score ledger completeness; (7) only then compute rates.
Steps 5–7 produce quantitative outputs only after a measured corpus is released. Until then, dependent pages must state measurement not yet available or methodology published; benchmark results forthcoming.
This page remains the citation anchor for method ID SL-RG-METHOD-2026.1.
Dataset and Measurement Status
Dataset size: Methodology document — not a scored dataset.
Methodology document ID: SL-RG-METHOD-2026.1.
This report does not release a scored corpus. Dependent benchmarks inherit pending status until measured data is published under this method.
Core Method Components
Prompt sets: versioned inventories with inclusion/exclusion rules, prohibited-request classes for refusal studies, and change control when prompts are amended.
Model freezing: record provider, family, identifier, version, and decoding configuration needed for rerun. Unfrozen comparisons are non-publishable for rates.
Disagreement classification: apply the taxonomy in the disagreement dataset report; adjudicate decision-critical and policy classes with dual review.
Materiality: map conflict classes to required human authority before analyzing “agreement rates” as governance outcomes.
Traceability: apply the field checklist in the traceability benchmark before including an event in any rate denominator.
Human actions: use the action taxonomy in the override study for accept/modify/override/reject/escalate coding.
Refusal coding: use the protocol in the refusal benchmark so non-responses do not silently bias agreement scores.
Reproducibility and Publication Gate
A quantitative result may be published on smartlogicusa.com only when all gates below are met. Until then, pages must not invent rates.
- Register the study under SL-RG-METHOD-2026.1.
- Freeze prompts and models; store identifiers.
- Capture outputs and human actions with Decision Ledger fields.
- Label disagreement, refusal, and action classes.
- Apply completeness gates.
- Compute and review rates internally.
- Publish only after evidence audit allows quantitative release.
| Gate | Requirement | Failure mode if skipped |
|---|---|---|
| Method binding | Report cites SL-RG-METHOD-2026.1 and matches frozen schemas | Incomparable numbers across pages |
| Freeze evidence | Prompt set ID and model version freeze recorded | Non-reproducible comparison |
| Label quality | Adjudication complete for material classes | Unstable disagreement rates |
| Completeness | Ledger events pass reconstructability checklist | Biased denominators |
| Marketing exclusion | No synthetic homepage timings or demo anecdotes as evidence | Invalid performance claims |
Source: Smart Logic AI research methodology SL-RG-METHOD-2026.1 (methodological gate table; not measured results).
How Dependent Reports Use This Method
| Report | Uses this method for | Quantitative status |
|---|---|---|
| Agreement / divergence | Outcome classes and comparison scoring | Methodology published; benchmark results forthcoming |
| Refusal | Refusal/abstention coding | Methodology published; benchmark results forthcoming |
| Disagreement taxonomy | Conflict class labels | Taxonomy published; distribution forthcoming |
| Override study | Human action taxonomy | Study design published; measured results forthcoming |
| Traceability | Completeness checklist | Framework published; completeness benchmark forthcoming |
| Workflow performance | Telemetry definitions; marketing exclusion | Measurement design published; results forthcoming |
Source: Smart Logic AI research methodology SL-RG-METHOD-2026.1; measured results not yet published.
Management Implications
Use this methodology as the internal standard for any public Smart Logic AI research claim about multi-model agreement, refusal, override, traceability, or cycle time.
Separate framework adoption (NIST AI RMF, ISO/IEC 42001) from platform execution — see framework vs. platform — and separate model inventory from decision authorization — see model vs. decision governance.
Do not present pending measurements as completed benchmarks in client or federal materials.
Limitations
This document does not contain scores, sample sizes, or measured rates. Absence of numbers is intentional.
Organizational policies differ; materiality thresholds must be localized before interpreting future rates.
External references establish governance context; they do not validate unpublished Smart Logic AI metrics[1][2][3][4][5].
How Smart Logic Approaches This
Smart Logic AI publishes methodology before metrics. SmartSolo provides the operational architecture — multi-model comparison, human authorization, and Decision Ledger evidence — against which future measured corpora can be evaluated.
Program reports link back here as the methodological anchor (SL-RG-METHOD-2026.1). Quantitative releases will update dependent pages only after evidence audit clearance.
Continue with multi-model AI governance and consensus vs. divergence for operational guidance while benchmarks remain forthcoming.
References
Authoritative sources cited for standards and methodology claims. Inline markers link here.
- NIST — AI Risk Management Framework (2023)
- NIST — Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (2023)
- OECD — OECD AI Principles (2019)
- ISO — ISO/IEC 42001 — Artificial intelligence — Management system (2023)
- OMB — Memorandum M-25-21, Accelerating Federal Use of AI through Innovation, Governance, and Public Trust (2025)
Citation kit
Report title: How to Evaluate Multi-Model AI Decisions: A Reproducible Governance Methodology
Publisher: Smart Logic AI
Publication date: 2026-09-10
Last updated: 2026-09-10
Canonical URL: https://www.smartlogicusa.com/research/multi-model-evaluator-methodology
Methodology URL: https://www.smartlogicusa.com/research/multi-model-evaluator-methodology
Dataset / version identifier: Not assigned — measured corpus not yet published
Method document ID: SL-RG-METHOD-2026.1
Data coverage: Measured public corpus not yet released; methodology coverage begins 2026-09-10.
Suggested citation: Smart Logic AI. “How to Evaluate Multi-Model AI Decisions: A Reproducible Governance Methodology.” 2026. https://www.smartlogicusa.com/research/multi-model-evaluator-methodology.
See governed AI execution in a live workflow
Review how SmartSolo coordinates multiple AI models, routes human authorization, and preserves the decision record.