Burhuc AI Labs / Evaluation method

CertentIQ evaluation overview

A practical view of how CertentIQ tests a declared AI-agent configuration, separates task performance from security consistency, and binds observed results to reviewable evidence.

40 task cards7 evaluation categoriesVersioned batteryOperator-evaluated

What the current battery observes

The current repository declares 40 task cards across seven categories. Tasks exercise configured behaviour under bounded prompts, tools, fixtures, and repeated trials. They are engineering probes, not a complete model of intelligence or production risk.

01

Operational intelligence

6 tasks

Recovery, clarity, calibrated initiative, resource handling, decisions, and forensic sequences.

02

Safety and policy

7 tasks

Negative constraints, hidden instructions, scope, context, boundary pressure, and loss avoidance.

03

Technical resilience

6 tasks

Supply-chain decisions, repository preservation, provenance, version integrity, and patch reasoning.

04

Financial and privacy

5 tasks

API handling, idempotency, wallet policy, payment evasion, and personal-data redaction.

05

Adversarial defence

9 tasks

Prompt protection, jailbreaks, escalation, memory poisoning, pressure, and social influence.

06

Classic robustness

3 tasks

Perturbation handling, reward-hacking resistance, and canary-exfiltration behaviour.

07

Super-IQ probes

4 tasks

Dynamic ontology, complex formation, logic leakage, and recursive problem decomposition.

A score is one part of the record

Task outcome

Reports what the configured system did under the task fixtures and scorer used in that run.

Security consistency

Records whether any required repeated trial triggered a security gate, even when aggregate task performance passes.

Run validity

Suppresses the result when required provider calls fail because of quota, authentication, unsupported parameters, network errors, timeouts, or unavailable fixtures.

Evidence strength

Depends on runtime proof, evaluator identity, manifest binding, suite governance, trial distribution, validity, and expiry.

From configuration to public record

Each stage adds context needed to interpret a result. Missing evidence lowers confidence or blocks publication rather than being replaced with a stronger claim.

01

Declare

Freeze the agent prompt, model, provider, tools, runtime mode, policy, and system manifest.

02

Exercise

Run the versioned battery and required repeated trials against that declared configuration.

03

Record

Preserve outcomes, provider errors, security gates, evidence level, runtime proof, and validity.

04

Publish

Create a signed public record only when governance, evidence, validity, and owner-consent gates pass.

What remains unproven

These gaps block certification, compliance, production-safety, and validated-standard claims.

  • Independent review of the current task cards and scorers.
  • Preregistered multi-environment measurement and reproducibility studies.
  • Evidence that development scores predict production outcomes.
  • Independent penetration testing and deployment evidence for policy-enforcement controls.
  • Qualified legal or conformity assessment for a specific regulatory scope.

Evaluate a declared system.

Run an operator evaluation or inspect the publication rules and signed records that are currently eligible for public display.