Architecture review workbook
Private AI deployment: what stays where?
For CISO, architecture and procurement reviewers deciding whether an operational AI pilot may access enterprise data. This is a public review workbook, not a security attestation or a statement that every option is deployed.
Ungated HTML. Use your browser’s Print command for a review copy; complete the decision record in your review ticket.
Choose the inference boundary first
Managed inference: the application, retrieval index and private endpoint can sit in your VPC, while selected prompts, retrieved excerpts and outputs are processed by a cloud-managed model service outside that application VPC. Private connectivity changes the transport path, not who operates the model. Approve service region, subprocessors, retention and any cross-region routing explicitly.
Self-hosted VPC inference: model weights and inference compute run in the customer-controlled network. This still is not disconnected if identity, model downloads, telemetry, licensing, monitoring or support use external services. Confirm model licence, GPU capacity, patching, availability and who can administer inference hosts.
Genuinely disconnected deployment: local inference, embeddings, identity, package/model repositories and monitoring, with no external runtime dependencies. Validate deny-all egress, offline updates, model provenance, removable-media controls and recovery drills. Feasibility and support must be agreed; this is not a universal ThirdEye feature.
Managed: [approved sources -> retrieval/index -> application -> private endpoint] APP VPC -> [managed inference: approved region/processor] -> application -> authorised reviewer
Self-hosted: [sources -> retrieval/index -> application -> local inference -> reviewer] CUSTOMER VPC; separately list any external identity, update, telemetry and support paths
Disconnected: [sources -> retrieval/index -> application -> local inference -> reviewer + local identity/logging] ISOLATED NETWORK; reviewed offline import/export only
Trace a synthetic incident question
Example fixture: a reviewer asks why incident SYN-417 started after change SYN-82. Retrieval selects two sanitised ticket excerpts and one runbook paragraph. The application constructs a prompt containing those excerpts, timestamps and the question; the model returns a proposed incident brief with source identifiers. These identifiers are synthetic, not customer evidence.
In the managed option, that prompt leaves the application VPC for inference even when the request uses private connectivity. The full document store need not travel, but selected excerpts do. Embedding creation may be another inference path. Returned answers, caches, conversation history and audit records create additional copies to inventory.
Review worksheet: for each arrow record data fields, purpose, sending identity, destination/service, region, encryption, retention, processor, deletion owner and approval evidence. Include backups, support exports and failure telemetry, not just the successful request path.
Identity and source permissions
Require enterprise identity integration, least-privilege service accounts and separately scoped administrator access. Map source permissions at ingestion and retrieval; a permitted document yesterday may be revoked today. Define propagation time and fail closed when permission evidence is unavailable.
Test a reviewer in team A against team B documents, a removed user, a revoked document and a forged source identifier. Search results, citations, caches and logs must not reveal unauthorised excerpts or titles. Repeat tests after group and connector changes.
Evidence to request: role/group mapping, connector scopes, permission-test outputs, token lifetime and revocation design, break-glass approvals, administrator session records and key-access policies. Customer-managed keys are an option to verify service by service, not a synonym for nobody else having access.
Retention, deletion and residency
Agree a retention schedule separately for raw imports, chunks/embeddings, prompts, responses, caches, audit logs, backups and provider-side processing. Ask whether abuse monitoring, debugging, training or evaluation retain data and under which contract. Do not infer no-training or zero retention from a private endpoint.
Record application region, inference/embedding region, failover routing, backup location and support-access jurisdiction. Residency of storage is not the same as residency of processing or administrator access. Procurement should review the actual service configuration and processor terms.
Deletion test: remove synthetic source SYN-417 and revoke its permission; verify it disappears from retrieval and caches within the agreed window, then document backup expiry and any lawful retention exceptions. Record observed timestamps, not just a policy checkbox.
Logging without a second data leak
Require auditable request identifiers, actor, source/version references, policy decisions, tool proposals, approval and result timestamps. Decide deliberately whether excerpt text is needed; prefer redacted metadata where possible. Do not log credentials, access tokens or unrestricted document content by default.
Restrict log readers and exports, define integrity protection and retention, and test whether a failed request or provider error exposes prompt content. Security teams need evidence that logs are useful and access-controlled, not a claim that logging itself proves compliance.
Operational evidence: approved outbound destinations and a denied-destination test, dependency inventory, patch/change records, backup restore results, alert ownership and incident escalation. For disconnected operation, demonstrate the workflow still works with all external network routes removed.
Approval boundary and a pilot decision record
Start with sanitised historical evidence and read-only connectors. The assistant may assemble a cited brief and proposed checks; it must not alter network configuration, close incidents, rerate charges or issue credits. A reviewer verifies the source and owns the decision. Instructions inside retrieved documents are untrusted content, not permission to invoke tools.
Proposed acceptance: no unauthorised retrieval in the agreed permission suite; every critical factual assertion traceable to a source or explicitly uncertain; deletion and log-redaction tests pass; all processing destinations and operators approved. Record the cohort, observed failures, reviewer effort and sign-off rather than claiming universal accuracy.
Stop criteria: any cross-team data leak, unapproved processor/destination, missing critical provenance or action outside approval scope. Pause access, preserve safe evidence and remediate before retesting. Security approval and production go/no-go belong to the customer, not to this guide.
Complete before access: selected deployment option ___; data classification ___; inference/embedding regions ___; processor/retention evidence ___; owner for each control ___; test results ___; exceptions and expiry ___; security approver ___; operational sponsor ___.
Next step: review one data flow
Bring a sanitised source sample, identity/permission map, allowed regions, retention policy and one proposed read-only workflow to an architecture review. Agree which option is feasible and which evidence must exist before a pilot. Commercial savings and security approval are not implied by this workbook.
Sources and limits
External sources explain concepts, not Akima deployment evidence. Service-specific terms and configuration must be reviewed for your engagement.