CertentIQ / Platform architecture

Technical Architecture - Development Release

A platform-level view of evaluation, remediation research, policy-mediated controls, shared services, and the evidence required to interpret each result without extending it beyond the tested system and version.

Development releaseOperator-evaluatedEvidence-boundExternal validation incomplete

What this architecture represents

This document describes the current development direction and implemented repository boundaries. It is not a normative technical standard, an assurance level, or a representation that every component is production-deployed.

Document typeDevelopment architecture

A maintained description of components, responsibilities and evidence limits.

Evaluation statusActive validation

Task, scorer and measurement review remains in progress.

Deployment statusComponent specific

Repository presence does not establish production operation.

Assurance statusNo external conclusion

Independent validation and conformity assessment are incomplete.

Evaluate, develop, then mediate

The three product areas serve different development purposes. A result from one stage does not automatically authorize promotion into the next.

01

CertentIQ

Runs versioned, adversarial and tool-enabled tests against a declared agent-system configuration. It records observed behavior, trial distributions, security-gate outcomes and run context.

Current evidenceOperator-evaluated run records. No independent measurement validation or general production prediction.
02

CertentiTrain

Uses observed failure patterns to generate remediation candidates and controlled practice work for the same declared configuration.

Current evidenceDevelopment workflow. Improvement requires protected re-evaluation and remains system, version and suite specific.
03

CertentiWall

Represents the policy-mediated control boundary intended to constrain selected effects and preserve decision evidence around deployed agent actions.

Current evidenceArchitecture and implementation work. Deployment effectiveness requires separate topology review, threat modeling and penetration testing.

Platform responsibilities and boundaries

Shared services support reproducibility, access control and record integrity. Their effectiveness remains dependent on configuration, deployment and complete mediation of protected actions.

ServiceImplemented responsibilityEvidence boundary

Identity and authorization

Signed tokens, issuer and audience checks, action scopes, session controls, revocation and protected routes.

Does not establish legal identity, complete access governance or production key custody.

Evaluation governance

Versioned suites, battery fingerprints, task and scorer bindings, repeated trials, provider-error handling and publication gates.

Does not establish construct validity, scorer correctness or cross-version comparability.

Tool policy

Bounded tool interfaces, path containment, egress policy, action telemetry and protected-effect checks.

Does not prove operating-system isolation, complete mediation or tenant separation.

Evidence integrity

System manifests, run context, signed records, public-key discovery, validity periods and revocation state.

Signature validity does not establish evaluator independence or behavioral validity.

Storage and publication

Private-by-default results, explicit publication consent, supported state backends and public-record contracts.

Does not establish retention compliance, disaster recovery or production resilience.

Three questions for every result

A score or control observation is interpretable only when the tested system, evaluation declaration and evidence level remain attached.

System

What exactly was tested?

Identify the model, provider, prompt and policy configuration, tools, runtime mode, manifest digest and material dependencies.

Evaluation

Under which test contract?

Identify the suite, task and fixture versions, scorer logic, trial policy, battery fingerprint, provider outcome and security gates.

Evidence

Who produced and reviewed it?

Distinguish operator-evaluated, independently reproduced and externally assessed evidence. Current public records must not imply a higher level than they carry.

Promotion requires fresh evidence

Architecture labels do not substitute for release criteria. Each component needs evidence appropriate to the change and intended use.

  • Record the exact code revision, dependencies, configuration, provider, model, fixtures, scorers and runtime-proof mode.
  • Invalidate or qualify comparisons after material changes to prompts, tools, policies, providers, tasks or scoring logic.
  • Keep aggregate performance separate from failed required security trials and unavailable evidence.
  • Require deployment-specific review for infrastructure, secrets, network, identity, monitoring and operational controls.
  • Use standards mappings as engineering references only; qualified assessment is required for any conformity conclusion.

Move from architecture to inspectable evidence.

Review the candidate evaluation architecture for the run model, or inspect the repository regression report for current secure-development test scope.