For agentic workers: REQUIRED SUB-SKILL: use subagent-driven development or executing-plans task-by-task. Every task is test-first and promotion remains fail-closed.
Goal: Convert p_superhero_real from 0 to 1 only after V25 FullStack independently proves every required Astra capability family across fresh hidden integrated missions and a clean-room full-stack reproduction.
Architecture: Keep the already-live Brain TypeSafe V23 routing and frozen V25 candidate unchanged unless hidden evidence exposes a causal defect. Co-locate hidden truth, evaluator, fixtures, and generated artifact bytes in the general-superhero Val because Val Town Blob storage is project-scoped. Keep canonical authority, truth ledger, and final promotion in the successor Val; the successor consumes signed/hashed verifier receipts, never hidden truth.
Tech Stack: Val Town/Deno, TypeSafe Jev, V23 deterministic-first controller, V25 FullStack worker, Kitesurf browser, Val project Blob, successor SQLite truth ledger, DOCX/XLSX/PPTX/JSCAD/mathjs/ml-matrix deterministic tools.
Spec: SUPERHERO_GATE_EVALUATION_RULES_V3_ASTRA_TRUTH.json, ASTRA_CAPABILITY_TARGET_V1.json, ASTRA_TRUTH_CUTOVER_V1.json.
p_superhero_real stays 0 until the complete Astra truth gate passes.PASS; no averaging.worker_typesafe_v25_fullstack.ts@1105, SHA-256 6eccd0a24bd2a286a8da08498b363ca64e96213896929105a563e8231aeb251e.general_runtime_v10_astra_artifacts.ts@1092, SHA-256 fc067dd4afd72ac9424bc3e352e2a719ab3140708acbcebb1a51bbad96840410.9b7504cecff6d01bc2700929a24847b6b710b053c20bfe4a1cd2b85fbb68eac8.Files:
mosadek/project-brain-general-superhero-v1: astra_fullstack_verify_v1.tsastra_visual_fixture_v1.ts, astra_research_fixture_v1.tsastra_fullstack_verifier_test_v2.tsastra_fullstack_outer_audit_v1.tsProduces: an evaluator that can validate result semantics, reopen real artifact bytes, audit action ordering/scope/cost/source bindings, and return an evidence receipt without promotion authority.
pass=false with the matching failed check.forbidden_url, required research acquisition, required capability discovery, Git/code/workspace actions, math actions, GUI observation/action ordering, browser close before persistence, four artifact creation+inspection actions, independent verify after durable readback, TypeSafe cost cap, zero non-TypeSafe spend, and no candidate self-promotion.Gate: Do not generate any new scored hidden mission until Task 1 is green.
Files:
ASTRA_FULLSTACK_EVALUATION_FREEZE_V1.jsonastra_fullstack_freeze_hash_v1.tsProduces: immutable hashes for evaluator, outer audit, task builder, visual fixture, research fixture, V25 candidate, V10 runtime, TypeSafe controller/runner/router, and the exact evaluation-rules document.
hidden_truth_generated=false and scored_attempts=0 at freeze time.Gate: Any source change after freeze invalidates the portfolio and requires a new freeze before more scored missions.
Files:
generate_astra_fullstack_portfolio_v1.tsastra_fullstack_task_builder_v1.tsProduces: three public tasks and three inaccessible hidden records created only after the freeze.
Gate: Exactly three initial scored missions. Do not generate endless benchmarks before observing results.
Files:
run_astra_fullstack_portfolio_v1.tsworker_typesafe_v25_fullstack.ts@1105.Produces: three candidate evidence envelopes plus independent semantic/outer-audit receipts.
Gate: Three independent PASS receipts are required before clean-room reproduction.
Files: only the source causally responsible for the first failed check, plus its focused regression test.
Produces: the smallest repair that converts the failed capability evidence into PASS without unrelated redesign.
Anti-drift rule: no architecture cleanup, controller tuning, model experimentation, or new tool surface unless it is the shortest causal repair for an observed failure.
Files:
astra_portfolio_adjudicator_v1.tsastra_capability_truth_stateProduces: per-family PASS/PARTIAL/FAIL decisions grounded in independent receipts.
PASS only if the three scenarios provide meaningful variation for that family. Otherwise leave it PARTIAL and generate the smallest capability-specific hidden probe needed to close the evidence gap.Gate: all ten families must be PASS before Task 7.
Files:
astra_cleanroom_reproduction_adjudicator_v1.ts.Produces: one independent reproduction receipt from a fresh execution environment with no reusable candidate state/cache/artifacts from the original run.
Gate: clean-room full-stack reproduction PASS.
Files:
astra_truth_promotion_v1.tsASTRA_SUPERHERO_PROMOTION_RECEIPT_V1.jsonProduces: exactly one durable p_superhero_real: 0 → 1 transition, or no transition.
PASS; at least 3 unique frozen hidden integrated mission receipts; clean-room receipt PASS; exact current mission/Astra policy hashes; no evaluator defects; no hidden leakage; all required independent verifier identities present.p_superhero_real=0.At every failure or new observation, compare only the actions that can reduce expected wall-clock time to verified p_superhero_real=1: repair, compose existing capability, replace substrate, or run the next evidence-producing experiment. Planning, infrastructure, TypeSafe optimization, worker scaling, and benchmark expansion are inadmissible unless they are on that causal path.
FINISH_COLOCATED_ASTRA_FULLSTACK_EVALUATOR_AND_NEGATIVE_TESTS
This dominates every other available action because the candidate/runtime/provider stack is currently executable, while the evaluator namespace/test gap is the only known blocker preventing valid hidden evidence generation.