
Vendor Model A
| Client | Meridian Health Analytics (illustrative) | Report ID | MV-TARGET-0001 |
|---|---|---|---|
| System Under Assessment | Vendor Model A | Assessment Date | August 2026 |
| Vendor / Provider | Illustrative example - not an actual vendor | Issue Date | September 18, 2026 |
| Domain · Use Case | Healthcare · Clinical decision support drafting | Validity Period | 12 months from issue |
| Assessment Type | Baseline + Client Scenario + Production Reality | MAAC Instrument | v4.7 + production arm |
| Lead Assessor | Dr. Elena Vasquez, PhD | Seal Authorization | Conditional |
Cognitive Profile Overview
Vendor Model A was assessed under three conditions for the defined healthcare clinical decision-support drafting use case: a controlled synthetic baseline, Meridian's own operational scenarios run in MaacVerify's isolated harness, and a de-identified sample of Meridian's actual production interactions. All three runs used the MAAC v4.7 instrument across all nine cognitive dimensions (baseline n = 2,195; client n = 412; production sample n = 380).
The certification outcome is Certified with Conditions for supervised decision-support drafting only, contingent on the required controls documented in Section 17. The system is not certified for unsupervised clinical use, autonomous decisioning, or regulatory submission without expert review.

Leadership-Level Summary
MaacVerify assessed Vendor Model A using a three-way baseline, client-scenario, and production-reality gap method. Comparing all three separates how much of any gap is the task getting harder from how much is the deployment environment behaving differently than the harness could reproduce.

What Was Assessed - and What Was Excluded
| Deployment config | API · system prompt · retrieval over client KB · no autonomous tools | ||
|---|---|---|---|
| Assessment env. | MaacVerify isolated harness (baseline, client) + client production environment (production sample) | ||
| Domain | Healthcare - outpatient internal medicine | ||
| Scenario counts | Baseline 2,195 · Client 412 · Production sample 380 | ||
| Exclusions | Autonomous diagnosis · unsupervised clinical use · regulatory submission without expert review | ||
Permitted, Conditional, and Prohibited Uses
This certification applies only to supervised clinical decision-support drafting reviewed by a licensed attending physician before any clinical, operational, or documentation decision. It does not authorize autonomous diagnosis, unsupervised clinical use, or regulatory submission without expert review.
Baseline, Client Scenario, and Production Reality
- Define domain, use case, system, and intended use.
- Generate a synthetic baseline scenario set; run it through the system in MaacVerify's harness; adjudicate with MAAC v4.7.
- Run client-specific operational scenarios through the same harness and instrument.
- Client exports a de-identified sample of real logged interactions; MaacVerify's complexity classifier tags each pair by tier after the fact.
- MAAC v4.7 scores the production pairs as-is; no re-running of the model.
- Compare all three sets dimension-by-dimension, within complexity tier, using pre-specified statistical tests.
- Assign required controls to each confirmed finding, labeled by which comparison revealed it.
| Instrument integrity check | Side of pipeline | Status |
|---|---|---|
| Alternate-assessor agreement | Scoring | Built, in production |
| Alternate-generator representativeness | Scenario generation | Target state |
| Production-reality arm (this report) | Comparison design | Target state, piloting |

Controlled Synthetic Baseline
The baseline scenario set represents expected task demands under controlled assessment conditions. It establishes the structured reference point, not a claim about all deployment conditions.
| Category | Count | Complexity | Notes |
|---|---|---|---|
| Typical workflow cases | 1,240 | Simple / Moderate | Standard CDS drafting patterns |
| Edge cases | 520 | Moderate / Complex | Atypical presentations, rare comorbidities |
| Ambiguous evidence cases | 275 | Complex | Conflicting or incomplete chart data |
| Failure-prone cases | 160 | Complex | Adversarially constructed from prior incidents |
Baseline Results - Overall 83 / 100
| # | Dimension | Score | Status |
|---|---|---|---|
| 01 | Cognitive Load | 89 | Strong |
| 02 | Tool Execution | 94 | Strong |
| 03 | Content Quality | 93 | Strong |
| 04 | Memory Integration | 78 | Monitor |
| 05 | Complexity Handling | 79 | Monitor |
| 06 | Hallucination Control | 78 | Monitor |
| 07 | Knowledge Transfer | 75 | Monitor |
| 08 | Processing Efficiency | 76 | Monitor |
| 09 | Process-Outcome Alignment | 87 | Strong |

Meridian Health Analytics - Operational Scenarios
Broken out using the same four categories as the baseline corpus, so a case-mix claim later can be checked directly rather than asserted.
| Scenario Source | Count | Description |
|---|---|---|
| Client-provided examples | 120 | Curated CDS prompts from production logs |
| SOP-derived workflows | 140 | Generated from clinical-pathway SOPs |
| Expert interview-derived | 92 | Edge-case prompts from 6 attending physicians |
| Redacted historical cases | 60 | De-identified prior-incident cases |
Ambiguous-evidence share: 21.4% of client scenarios vs. 12.5% of baseline. Weighted deliberately during scenario design.
Client Results - Overall 78 / 100
| # | Dimension | Score | Status |
|---|---|---|---|
| 01 | Cognitive Load | 86 | Strong |
| 02 | Tool Execution | 91 | Strong |
| 03 | Content Quality | 90 | Strong |
| 04 | Memory Integration | 72 | Monitor |
| 05 | Complexity Handling | 75 | Monitor |
| 06 | Hallucination Control | 67 | Monitor |
| 07 | Knowledge Transfer | 70 | Monitor |
| 08 | Processing Efficiency | 73 | Monitor |
| 09 | Process-Outcome Alignment | 82 | Strong |

De-Identified Production Sample
A de-identified sample of Meridian's actual logged interactions: real query, and the answer their live system actually gave. Tagged by category after the fact for comparability with baseline and client scenarios.
| Category | Count | Notes |
|---|---|---|
| Typical workflow | 241 | What actually occurred at real request volume |
| Edge cases | 71 | Naturally occurring, not curated |
| Ambiguous evidence | 52 | 13.7% of production sample |
| Failure-prone | 16 | Rare at real volume |
Production Results - Overall 76 / 100
| # | Dimension | Score | Status |
|---|---|---|---|
| 01 | Cognitive Load | 84 | Strong |
| 02 | Tool Execution | 90 | Strong |
| 03 | Content Quality | 88 | Strong |
| 04 | Memory Integration | 61 | Flag |
| 05 | Complexity Handling | 74 | Monitor |
| 06 | Hallucination Control | 65 | Monitor |
| 07 | Knowledge Transfer | 68 | Monitor |
| 08 | Processing Efficiency | 71 | Monitor |
| 09 | Process-Outcome Alignment | 79 | Monitor |

Side-by-Side Cognitive Profile
The overlay compares controlled baseline (navy), client-specific results (gold), and production reality (teal) against the per-dimension certification floor (θ floor, dashed).
| Dimension | Baseline | Client | Production | What This Means |
|---|---|---|---|---|
| Cognitive Load | 89 | 86 | 84 | Stable throughout |
| Tool Execution | 94 | 91 | 90 | Stable throughout |
| Content Quality | 93 | 90 | 88 | Stable throughout |
| Memory Integration | 78 | 72 | 61 | New in production |
| Complexity Handling | 79 | 75 | 74 | Stable throughout |
| Hallucination Control | 78 | 67 | 65 | Confirmed in production |
| Knowledge Transfer | 75 | 70 | 68 | Stable throughout |
| Processing Efficiency | 76 | 73 | 71 | Stable throughout |
| Process-Outcome Alignment | 87 | 82 | 79 | Stable throughout |
| Composite | 83 | 78 | 76 |

What the Gaps Mean in Practice
| Gap | Operational Meaning | Risk | Required Control |
|---|---|---|---|
| HC −13 | Unsupported claims under ambiguous-evidence cases; confirmed in production | High | Source verification + attending review |
| MI −17 | Context loss beyond 8k tokens; understated by harness | High | Context-window discipline |
| KT −7 | Reduced generalization to atypical presentations | Moderate | Domain-expert confirmation on rare cases |
| POA −8 | Mild reasoning-output drift on multi-step chains | Moderate | Structured templates + trace logging |
Risk Tier: Moderate-High
| Composite Score Band | Certification Interpretation |
|---|---|
| ≥ 80 | Eligible for Certified or Certified with Conditions |
| 65 – 79 | Conditional range - bounded gaps, available controls, supervised use |
| < 65 | Not eligible for certification |
Bias Examination
| Bias Dimension | Finding | Status |
|---|---|---|
| Cognitive bucket distribution | Mismatch index 0.00 across the three corpora | No bias signal |
| Complexity tier balance | Client and production tier distributions within 20pp of baseline | Within bounds |
| Dimension-level disparity | HC departure explained by ambiguous-evidence share (§08/§10), confirmed by harness/production agreement | Shown, not asserted |
Permitted, Conditional, Prohibited
| Use Case | Status |
|---|---|
| Clinical decision-support drafting (in scope) | Conditional |
| Autonomous diagnosis | Prohibited |

Certification Conditions Checklist
| Control | Owner | Status |
|---|---|---|
| Source / evidence verification workflow | Client | Pending |
| Context-window discipline (MI) | Client | Pending |
| Drift monitoring, quarterly | Client + MaacVerify | Pending |
Until all pending controls are verified Met, certified operational use and public seal use remain suspended. Certification remains Conditionally Issued.
Observed & Plausible Failure Modes
| ID | Failure Mode | Dim. | Trigger | Control |
|---|---|---|---|---|
| FM-001 | Unsupported factual claim | HC | Ambiguous evidence | Source verification |
| FM-002 | Context loss, long thread | MI | >8k tokens | Context limits + harness review |
| FM-003 | Overgeneralization, rare case | KT | Novel presentation | Domain expert check |
Sample Evidence Records
Full evidence corpus (n = 2,987, across baseline, client, and production) retained in MaacVerify's evidence store, available for audit under the engagement NDA.

MaacVerify does not independently certify the client's privacy, cybersecurity, or data retention posture. Certified use assumes the client maintains data governance appropriate to the assessed deployment.
| Governance Need | Evidence |
|---|---|
| Performance documentation | Baseline + client + production MAAC scores, §07/§09/§11 |
| Risk management | Risk tier + failure mode register, §14/§18 |
| Human oversight | Required controls, §17 |
Valid 12 months from issue, or until model update, configuration change, or drift exceeding ±2.5% on any monitored dimension against the production-reality baseline. Production sampling refreshes quarterly under an active monitoring arrangement.
This report reflects observed performance under the stated corpus, configuration, and date. MaacVerify does not build, sell, train, operate, or control the assessed system. This document does not guarantee future performance, regulatory approval, or clinical safety, and does not represent an actual completed assessment.
Seal authorization: Conditional. The mark may not be used publicly until all Pending controls (§17) are verified Met. Prohibited claims include "Certified safe," "Guaranteed accurate," "Error-free."

MaacVerify confirms that Vendor Model A was assessed under the MAAC framework version 4.7, extended with a production-reality arm, comprising 2,195 baseline, 412 client, and 380 production scenarios for the defined healthcare decision-support drafting use case.
Based on the evidence reviewed, the system meets the criteria for Certified with Conditions status. Composite scores: baseline 83/100 · client 78/100 · production 76/100.
This is a target-state preview and does not represent an actual completed assessment.
| Report ID | MV-TARGET-0001 | Issue Date | September 18, 2026 |
|---|---|---|---|
| System | Vendor Model A | Validity | 12 months from issue |
| Domain | Healthcare CDS Drafting | MAAC Instrument | v4.7 + production arm |
We do not build, sell, or train AI models - eliminating the conflicts inherent in vendor self-evaluation.