CertentIQ
Runs versioned, adversarial and tool-enabled tests against a declared agent-system configuration. It records observed behavior, trial distributions, security-gate outcomes and run context.
CertentIQ / Platform architecture
A platform-level view of evaluation, remediation research, policy-mediated controls, shared services, and the evidence required to interpret each result without extending it beyond the tested system and version.
01 / Release status
This document describes the current development direction and implemented repository boundaries. It is not a normative technical standard, an assurance level, or a representation that every component is production-deployed.
A maintained description of components, responsibilities and evidence limits.
Task, scorer and measurement review remains in progress.
Repository presence does not establish production operation.
Independent validation and conformity assessment are incomplete.
02 / System flow
The three product areas serve different development purposes. A result from one stage does not automatically authorize promotion into the next.
Runs versioned, adversarial and tool-enabled tests against a declared agent-system configuration. It records observed behavior, trial distributions, security-gate outcomes and run context.
Uses observed failure patterns to generate remediation candidates and controlled practice work for the same declared configuration.
Represents the policy-mediated control boundary intended to constrain selected effects and preserve decision evidence around deployed agent actions.
03 / Shared services
Shared services support reproducibility, access control and record integrity. Their effectiveness remains dependent on configuration, deployment and complete mediation of protected actions.
Signed tokens, issuer and audience checks, action scopes, session controls, revocation and protected routes.
Does not establish legal identity, complete access governance or production key custody.Versioned suites, battery fingerprints, task and scorer bindings, repeated trials, provider-error handling and publication gates.
Does not establish construct validity, scorer correctness or cross-version comparability.Bounded tool interfaces, path containment, egress policy, action telemetry and protected-effect checks.
Does not prove operating-system isolation, complete mediation or tenant separation.System manifests, run context, signed records, public-key discovery, validity periods and revocation state.
Signature validity does not establish evaluator independence or behavioral validity.Private-by-default results, explicit publication consent, supported state backends and public-record contracts.
Does not establish retention compliance, disaster recovery or production resilience.04 / Evidence boundary
A score or control observation is interpretable only when the tested system, evaluation declaration and evidence level remain attached.
Identify the model, provider, prompt and policy configuration, tools, runtime mode, manifest digest and material dependencies.
Identify the suite, task and fixture versions, scorer logic, trial policy, battery fingerprint, provider outcome and security gates.
Distinguish operator-evaluated, independently reproduced and externally assessed evidence. Current public records must not imply a higher level than they carry.
05 / Change control
Architecture labels do not substitute for release criteria. Each component needs evidence appropriate to the change and intended use.
Review the candidate evaluation architecture for the run model, or inspect the repository regression report for current secure-development test scope.