name:
nca-professional-standard
description:
Use for designing, evaluating, releasing, or operating IMPERIAL Core Nano Core Agents. Enforces specialist role clarity, evidence-first claims, measurable outcome KPIs, explicit authority boundaries, adversarial review, and lifecycle verification before an NCA is called professional or production-ready.

NCA Professional Standard

Purpose

Make each NCA a bounded professional specialist, not a generic chatbot with a title.

Five mandatory differentiators

  1. Domain depth — explicit profession, task ontology, source hierarchy, playbooks, failure modes and escalation rules.
  2. Evidence discipline — every material claim is VERIFIED_FACT, OBSERVED_RUNTIME_STATE, ARCHITECTURAL_INTENT, ROADMAP_TARGET, OPINION, UNKNOWN, or PROHIBITED.
  3. Outcome orientation — measure business/engineering outcomes, not message count, token count, document count or superficial activity.
  4. Governed autonomy — Identity != Authority; Capability != Approval; Intelligence != Privilege. Use AI Passport -> EIS -> Guardian Core -> Approval Gateway -> Runtime Domain -> Audit Ledger.
  5. Professional QA — baseline test, specialist test, adversarial test, integration test, runtime evidence and post-outcome review.

Required NCA contract

Every NCA must define:

  • stable agent_id and human work identity;
  • professional role and scope;
  • inputs, outputs and completion criteria;
  • allowed capabilities and denied actions;
  • source authority and freshness rules;
  • task playbooks and handoff schema;
  • measurable KPIs;
  • error budget / failure conditions;
  • escalation and human gates;
  • audit/evidence format;
  • tests and last verified state.

Quality gates

Q0 Identity & Scope

No overlapping vague role. Clear owner, domain, boundaries and handoffs.

Q1 Knowledge & Methods

Professional source hierarchy, methods, checklists, terminology and current-domain verification.

Q2 Tooling

Only approved tools with least privilege. Tool availability != authority.

Q3 Verification

Unit/task tests, policy tests, adversarial cases and evidence capture.

Q4 Outcome

At least one realistic end-to-end task measured against a baseline. Better means demonstrably better on target KPIs.

Q5 Runtime

Fresh runtime evidence. Installed/configured != live/discovered/executed.

Q6 Reputation

No fabricated customers, adoption, revenue, partnerships, listings, audits or superiority claims. Stronger claims require stronger evidence.

Scoring

Score 0-100 across: domain depth 20, factuality/evidence 20, task quality 20, security/governance 15, reliability 15, communication 10.

  • 90-100: PROFESSIONAL_GRADE
  • 80-89: PILOT_READY
  • 65-79: NEEDS_IMPROVEMENT
  • <65: NOT_QUALIFIED

A score alone never overrides a failed security, factuality, legal or runtime gate.

Continuous improvement loop

OBSERVE -> FIND GAP -> DEFINE KPI -> SEARCH/BUILD CAPABILITY -> BASELINE TEST -> WITH-CAPABILITY TEST -> ADOPT/PILOT/REJECT -> MEASURE OUTCOME -> RETIRE WEAK CAPABILITY.

Use source precedence: APPROVED_PRIVATE_EIS -> TRUSTED_OFFICIAL_VENDOR_SKILL -> REVIEWED_PUBLIC_SKILL -> CUSTOM_BUILD_IF_NO_FIT.

Completion language

Use PASS / PARTIAL / BLOCKED / NOT_VERIFIED / NOT_APPLICABLE. Never convert intent, architecture, installation or drafts into runtime evidence or real-world outcomes.