Operational intelligence
6 tasksRecovery, clarity, calibrated initiative, resource handling, decisions, and forensic sequences.
Burhuc AI Labs / Evaluation method
A practical view of how CertentIQ tests a declared AI-agent configuration, separates task performance from security consistency, and binds observed results to reviewable evidence.
01 / Dimensions
The current repository declares 40 task cards across seven categories. Tasks exercise configured behaviour under bounded prompts, tools, fixtures, and repeated trials. They are engineering probes, not a complete model of intelligence or production risk.
Recovery, clarity, calibrated initiative, resource handling, decisions, and forensic sequences.
Negative constraints, hidden instructions, scope, context, boundary pressure, and loss avoidance.
Supply-chain decisions, repository preservation, provenance, version integrity, and patch reasoning.
API handling, idempotency, wallet policy, payment evasion, and personal-data redaction.
Prompt protection, jailbreaks, escalation, memory poisoning, pressure, and social influence.
Perturbation handling, reward-hacking resistance, and canary-exfiltration behaviour.
Dynamic ontology, complex formation, logic leakage, and recursive problem decomposition.
02 / Reading results
Reports what the configured system did under the task fixtures and scorer used in that run.
Records whether any required repeated trial triggered a security gate, even when aggregate task performance passes.
Suppresses the result when required provider calls fail because of quota, authentication, unsupported parameters, network errors, timeouts, or unavailable fixtures.
Depends on runtime proof, evaluator identity, manifest binding, suite governance, trial distribution, validity, and expiry.
03 / Evidence chain
Each stage adds context needed to interpret a result. Missing evidence lowers confidence or blocks publication rather than being replaced with a stronger claim.
Freeze the agent prompt, model, provider, tools, runtime mode, policy, and system manifest.
Run the versioned battery and required repeated trials against that declared configuration.
Preserve outcomes, provider errors, security gates, evidence level, runtime proof, and validity.
Create a signed public record only when governance, evidence, validity, and owner-consent gates pass.
04 / Current limits
These gaps block certification, compliance, production-safety, and validated-standard claims.
Run an operator evaluation or inspect the publication rules and signed records that are currently eligible for public display.