Skip to Content
Welcome to the Novantra documentation.
GuidesGovernanceModulesAI Quality Evidence

AI Quality Evidence

The AI Quality Evidence module records scoped evidence about AI-assisted behavior. It is for benchmark suites, suite versions, metric thresholds, evaluation runs, item-level results, reviewer adjudication, and acceptance gates.

Use it when a tenant needs to prove how a configured AI profile, document-treatment path, OCR path, retrieval setup, or fallback mode performed for a defined scope. It does not make global accuracy, legal-validity, certification, or service-level claims.

What It Owns

AI Quality Evidence owns:

  • benchmark suites for a selected capability, source pack, language scope, profile, or deployment mode;
  • benchmark suite versions that freeze the selected scope, dataset policy, metric policy, and evidence policy;
  • metric definitions and tenant/customer threshold profiles;
  • evaluation runs and item-level result records with safe refs and sanitized snapshots;
  • human adjudication decisions for accepted, rejected, overridden, or inconclusive evidence;
  • acceptance gate decisions over a scoped evaluation run and threshold profile;
  • audit export fields, compliance-posture signals, activity entries, and work item descriptors for review follow-up.

Do not paste raw document body text, prompts, provider responses, embeddings, vectors, confidential source text, or provider payloads into quality snapshots. Store scoped refs, counts, scores, hashes, warning keys, and reviewer decisions instead.

When To Configure It

Configure AI Quality Evidence before presenting or accepting a governed AI workflow that needs measurable readiness proof.

Typical scopes include:

  • document intelligence extraction quality;
  • source-grounded citation coverage;
  • unsupported claim rate;
  • OCR and layout quality;
  • reviewer acceptance;
  • bilingual parity and terminology consistency;
  • prompt-injection resistance;
  • latency and throughput in a specific deployment profile.

Use configured only when a suite, threshold, run, and review path exist for the selected scope. Use evidence-required when the product can support the topic but the tenant has not recorded proof yet.

Operating Model

Start with a benchmark suite for one product capability and one evidence scope. Add a suite version when the scope, dataset policy, metric policy, or evidence policy needs to be frozen for review. Define metric thresholds that reflect the tenant’s acceptance criteria. Record evaluation runs from controlled sources and attach item-level results where the tenant needs traceability.

Reviewer adjudication is the controlled review layer. Use it when a run is inconclusive, when a threshold failed but a human override is justified, or when individual evidence needs reviewer context.

Acceptance gates are the final scoped decision layer over an evaluation run. Record a gate when the tenant has compared the run against the selected threshold profile and wants to mark the scope as accepted, rejected, inconclusive, or evidence-required. A gate decision can update the run status, but it still remains scoped to the suite, dataset, profile, language, deployment mode, and date captured in the snapshots.

For no-AI deployments, use the same evidence model to record deterministic, manual, fallback, or unavailable states. Do not create a separate no-AI evidence store.

Audit And Posture

Audit packages and posture views should use the module’s descriptors for suite key, suite status, metric family, run status, adjudication decision, acceptance gate decision, and acceptance threshold profile. These fields are safe references and statuses, not raw quality payload exports.

Work-management consumers can subscribe to adjudication-review and acceptance-gate-review descriptors when quality evidence needs human review. The AI Quality Evidence module does not own a separate task system.

Last updated on