Document Intelligence
Document Intelligence is the governed AI workflow for documents stored in the DMS. It helps teams summarize, review, classify, and extract structured candidates from contracts, policies, procedures, memoranda, notices, evidence packs, and scanned records.
Document Intelligence does not create a second document store or a separate governance register. The DMS remains the source of truth for documents and versions. AI Governance owns profiles, provider readiness, use authorization, action policies, runs, suggestions, OCR status, and evidence. Owning modules still apply final risks, findings, controls, obligations, evidence, or document metadata.
Content Extraction is the shared preparation layer for source-owned records. Document Intelligence consumes safe source refs, snippets, anchors, preview/OCR/search posture, and extraction readiness from that layer when available; it does not own extraction truth, OCR truth, preview truth, or vector storage.
When you would reach for this
Use Document Intelligence when:
- A reviewer needs a source-grounded summary of a long document.
- A team needs proposed parties, dates, deadlines, clauses, obligations, controls, risks, evidence references, policy requirements, or action items.
- A scanned PDF or image-based document needs OCR-aware handling before review.
- A document should be checked against approved tenant source packs, policies, standards, or glossaries.
- A regulated workflow needs evidence of the AI run, sources used, reviewer decision, and applied outcome.
Do not use it as a replacement for human review, document ownership, review-approval, change-management, or governed module workflows. The AI produces suggestions; Novantra records what happened; authorized people and owning modules decide what becomes final.
Before a document review
An administrator or AI governance owner should confirm:
- The Document Intelligence AI system is registered and active.
- The user has a valid AI use authorization for document extraction or review.
- An active Document Intelligence profile is published for the intended workflow.
- The profile is bound to an active provider connection and an approved model with the required capability for structured document output.
- The provider, model, and profile data-boundary settings allow the selected source classification and handling posture.
- The profile uses approved source packs, glossaries, and retrieval scope.
- An action policy is selected, usually
suggest-onlyorrequires-approvalfor regulated document work. - OCR is configured if scanned PDFs or images are expected.
- Metadata profiles or target appliers exist for any values that reviewers should apply.
Source packs are tenant-approved content. They can include internal policies, standards mappings, glossaries, templates, or licensed materials. Novantra does not require framework or regulatory body text to be baked into the product.
For a from-scratch setup, first add an AI provider in AI > Providers and approve the models that Document Intelligence may use. Then open AI > Document intelligence and follow the guided path:
- Create a profile, choose or create the governed AI capability, select the provider and approved model, describe the purpose, and define default extraction fields such as
summary,license number, orexpiry date. - Authorize the AI use for document metadata extraction, selecting the governed purpose, action mode, and allowed classifications.
- Configure a review policy for the profile. Human review is the usual default for regulated document work.
- Run extraction by pasting text or choosing an active DMS document, then review the generated suggestions before anything is applied.
Each default extraction field is a tenant-owned recipe entry. Administrators can add alternative labels, extraction guidance, expected formatting, missing-value handling, required-for-review posture, and confidence guidance. This lets a tenant tune how its approved AI model recognizes values such as trade-license expiry, legal name, clauses, or evidence references without editing raw prompts or provider payloads.
Novantra uses this guided setup to prepare the profile, authorization, review policy, and run configuration. Administrators only need advanced settings for specialized tenant policies. When a profile recipe is edited, publish a new profile version before running a test that should use the updated recipe.
The built-in AI starter pack imports a tenant-owned Document Intelligence readiness profile. It declares safe synthetic corpus categories, bilingual support, broad extraction targets, and review/privacy warnings. It is a starting point for configuration, not protected framework or regulator content.
The panel also reads profile and profile-version policy snapshots for operational metadata such as profile pack, source pack keys, supported languages, glossary key, evaluation suite key, manual attestation requirements, privacy review, redaction review, export-control review, and retention policy key. Missing declarations are shown as warnings. When a profile declares source packs or supported languages, reviewers choose the packs and language keys that define the run scope; the run stores those selected keys, not protected source-pack text. This does not mean the workflow is unusable, but it tells reviewers and operators which readiness evidence still needs tenant configuration.
When selected source-pack keys match existing AI context sources and chunks, the run also records sanitized source-pack context evidence: matched source counts, active/indexed status, chunk counts, language keys, rights/classification posture, and missing-context warnings. This is coverage evidence only. It does not copy source-pack body text, context chunk text, protected framework text, prompts, or provider output into the run snapshot.
When a source owner supplies source-graph evidence, the run records sanitized graph readiness evidence such as graph/scope/version refs, graph ref counts, node/edge/citation counts, coverage, drift status, and safe graph ref IDs. It does not copy graph payloads, raw source or OCR text, provider payloads, prompts, embeddings, vectors, storage keys, or source-owner internals into Document Intelligence evidence.
When a review uses a stored source artifact, Novantra resolves that artifact only through an active owner link that permits AI extraction. If the run names a source version, the artifact link must match that version before Document Intelligence can use it. Run evidence keeps safe artifact, source-version, extractor, checksum, rights, and citation references. It does not expose raw source bytes, storage keys, object paths, or raw extracted text in audit, posture, or export evidence.
Running from a DMS document
The primary workflow starts inside the document itself:
- Open a DMS document detail page.
- Open Content extraction when you need to inspect extraction, preview, keyword-search, OCR, or AI handoff readiness for the selected version.
- Open the Document intelligence right-lane tab or action.
- Review the readiness badges for profile, action policy, OCR, retrieval, and target appliers.
- Select the profile, the declared source packs and languages that should define the run scope, and the action policy.
- Optionally choose another document version in Compare with version, or paste an External comparison draft, when the workflow is a comparison or redline review.
- Review the profile recipe fields selected for the run. Advanced users can override the requested fields for a one-off test, but the normal path uses the profile recipe.
- Provide a reason and run the review.
- Load the run suggestions, inspect confidence, status, target, and proposed values.
- Request formal review for a suggestion when the workflow needs review-approval evidence, or approve/reject it directly when the reviewer has authority.
- Apply approved metadata values to the document metadata profile, or apply other reviewed suggestions through a registered owning-module target applier.
The panel also shows document security context when available, including classification, PII, and PHI indicators. Provider policy, model capability, model data-classification policy, data-boundary policy, and retention rules are enforced by the AI run and surrounding AI Governance controls. If the selected provider or model is unavailable, not approved, lacks the required capability, is not allowed for the source classification, or the data-boundary policy blocks the handoff, the run fails closed with sanitized evidence and no suggestions are created.
When the selected document or profile indicates sensitive handling, the run records those warning keys in the Document Intelligence readiness snapshot. Reviewers should treat warnings such as missing evaluation suite, missing glossary, sensitive source material, redaction review required, export-control review required, or manual attestation required as prompts for human validation before applying suggestions.
Quality metrics and manual attestation
Each completed Document Intelligence run records calculated quality posture from the run and suggestion evidence. The DMS panel shows suggestion count, average confidence, low-confidence count, missing values, citation-backed suggestions, source-graph ref coverage when supplied by a source owner, review outcomes, evaluation-suite declaration, OCR engine/page-anchor evidence when OCR contributed to the source text, and warning keys.
These metrics are calculated from sanitized run evidence. They do not expose raw prompts, provider responses, source document text, or extracted values in the quality summary.
When the selected profile requires manual attestation, when an evaluation suite is missing, when privacy, redaction, export-control, or sensitive-source warnings apply, or when a suggestion has low confidence or no citations, the reviewer remains in control. If the reviewer approves or edits a suggestion, the review reason is also stored as manual attestation evidence on that reviewed suggestion. Suggestions with those high-warning conditions cannot be applied through metadata or target appliers until the reviewed suggestion carries that attestation evidence. This gives auditors a human basis for the decision without creating a separate Document Intelligence approval ledger.
AI Governance can also materialize calculated monthly Document Intelligence portfolio metrics from these run snapshots, including average confidence, low-confidence suggestion totals, runs requiring manual attestation, and runs missing evaluation-suite coverage. Tenants can keep those calculated records, replace them with manual or attested records, or leave them as review prompts according to their operating model.
Bilingual and glossary scope
Profiles can declare supported language keys and a glossary or terminology pack. The DMS panel lets reviewers select the declared language keys for a run. The selected keys are validated against the profile, recorded in prompt/input evidence, and passed to configured providers as run scope.
Language selection does not create a Novantra-owned translation registry and does not override provider capability. If a profile does not declare supported languages or a glossary, the panel and run evidence show that as readiness warning material for the reviewer.
Readiness Evidence Checklist
Successful Document Intelligence runs also record a sanitized readiness evidence snapshot. The DMS panel turns that snapshot into a checklist covering DMS source context, profile pack declaration, source-pack scope, safe synthetic corpus guidance, citation coverage, source-graph refs, generated suggestions, comparison source evidence, OCR evidence, evaluation-suite declaration, privacy posture, and manual attestation posture.
The checklist is an operator aid. It helps a presenter see whether the selected run is strong enough for a regulated-enterprise demonstration, and which warnings still need explanation or reviewer action. It does not contain raw source text, prompt text, provider responses, or extracted values.
After suggestions are loaded, the panel adds live review/application status to the checklist. Review decisions and owning-module application are therefore based on the current suggestion statuses, while the static source/profile/OCR/quality evidence remains on the AI run.
OCR behavior
OCR is a governed AI/OCR capability selected for the deployment and tenant policy. The Document Intelligence panel shows OCR readiness:
- Configured means an OCR engine is available for documents that need extraction from images or scanned pages.
- Not configured means reviewers can still run text-based workflows, but image-only documents may not produce useful source text.
- Misconfigured or blocked means the workflow should be corrected before relying on OCR output.
Where OCR runs, run evidence includes engine or extractor information, normalized confidence where available, source-artifact references, page anchors, and region-anchor counts where policy permits. Low-confidence OCR output appears as a quality warning and should be treated as a review prompt, not as final truth.
Suggestions and application
Suggestions are reviewable AI outputs. Common targets include:
- Document metadata fields, applied through the metadata profile and field selected by the reviewer.
- DMS successor draft versions for reviewed proposed-version suggestions. The applier creates a new editable DMS draft version and records the reviewed AI value in the version change summary; it does not write the document body or approve the version.
- Governed findings, risks, controls, obligations, and evidence requirements, applied through registered target appliers.
- Privacy-review suggestions can create Privacy Processing DPIA records; export-control review suggestions can create draft Privacy Processing transfer records.
- Retention review suggestions can create Document Governance requirements, while explicit retention-policy or retention-rule suggestions can create draft Governed Retention rules.
- Change-management requests for proposed edits, redlines, or version-change candidates.
Redaction and masking suggestions remain governed finding or review-evidence candidates in this closure. They do not redact source documents, generate redacted renditions, or authorize external transfer by themselves; reviewers still use tenant-approved redaction, release, export-control, privacy, and retention workflows. If no target applier exists, the suggestion remains review evidence and cannot silently write to another module. This is intentional: AI does not bypass the module that owns final truth.
Formal review requests use the Review Approval foundation and target the AI suggestion record. They do not create a separate Document Intelligence approval ledger and they do not automatically apply the suggestion. After review is complete, an authorized reviewer still approves or rejects the AI suggestion before any owning-module application occurs.
Comparison workflows
When a DMS document has more than one version, the Document Intelligence panel can compare the selected version with another DMS version. The run records both source version references, marks the AI run as a comparison workflow, and produces reviewable suggestions for change summaries, redline-style notes, obligation deltas, and missing-reference checks.
The panel can also compare the selected DMS document with pasted external draft text. This is useful during a negotiation or policy review when a reviewer has a draft from outside DMS but does not want to import it as a governed DMS version yet. External draft text is transient run input: execution can use it for comparison, while run evidence stores only the draft name, text length, source kind, and sanitized content-policy metadata.
Version-change and redline suggestions route to Change Management target appliers where configured. Proposed-version suggestions can create a DMS successor draft version after review, but they still do not rewrite the document body or approve the version. Material document changes continue through DMS authoring, review, and change-management policy.
External draft comparison is an internal product workflow, not a public Document Intelligence API. Comparing against tenant source packs or catalog packs still depends on tenant-managed source-pack lifecycle and remains separate from the pasted-draft comparison flow.
Review, audit, and evidence
Every governed run records the profile, action policy, source references, warnings, suggestions, reviewer decisions, and action applications according to workspace retention policy. Reviewers must provide reasons for runs, approvals, rejections, and applications.
For audits, use Runs & Evidence to show:
- which document and version were reviewed;
- which profile, policy, and source scope were used;
- which profile pack, source packs, languages, glossary, evaluation suite, privacy warnings, and manual attestation requirements were declared;
- what calculated quality metrics and attestation warnings were recorded;
- what suggestions were produced;
- who requested formal review, approved, rejected, or applied each suggestion;
- which owning module or metadata field received the final applied value.
Regulated-Enterprise Readiness Path
A credible demo should use safe synthetic documents plus tenant-approved source packs:
- Import or configure the AI starter pack, then open a synthetic contract, policy, procedure, memorandum, notice, or evidence pack in DMS.
- Show profile, action policy, retrieval, OCR, and target-applier readiness.
- Run a source-grounded review.
- Compare two document versions and show version-change or redline suggestions routed to Change Management.
- Extract obligations, risks, controls, evidence references, parties, dates, deadlines, clauses, and action items.
- Show the quality and attestation posture for the run, including low-confidence or missing-evaluation warnings when applicable.
- Show the readiness evidence checklist and explain any warnings that remain.
- Request formal review for one material suggestion, then approve or reject suggestions with reasons.
- Apply at least one metadata suggestion to the document and, where configured, one DMS successor draft, governed finding, risk, control, obligation, evidence requirement, change-management request, privacy DPIA, privacy transfer record, retention rule, or document-governance requirement through a target applier.
- Open the run evidence to show source coverage, quality metrics, reviewer decisions, attestation evidence, and applied outcome.
For bilingual demos, use profiles and providers that explicitly support the language pair and a glossary/source pack approved by the tenant. If the selected provider or profile does not support the language, treat that as a visible warning and keep the reviewer in control.