Skip to content

Phase D Playbook — Question-Bank Mashup Cleanup (ongoing)

Portal doc. This is the master doc every Phase D session reads at start. Self-contained kickoff with state, plan, workflow, discovery infrastructure, and self-learning protocol. History detail is in the linked session handoffs (read only if needed).

Last updated: 2026-06-17 NZT session a40ade84 — 🎉 P3 FACTUAL SWEEP COMPLETE (96/96 batches, 1564 claims). factual_verify.py LLM sweep finished: VERIFIED 933 (59.7%) · STALE 132 (8.4%) · SCENARIO_FICTION 317 · UNVERIFIABLE 182. Triage queue: guided/files/factual-verdict.md. 8.4% real-drift yield vs 0.69% pricing-sweep baseline. Method pivot mid-run (Rule #5): research agents ran ~75 min/batch; switched batches 70-96 to general-purpose agents that WRITE their own verdict file (Sonnet, ~5 min/batch, parallel web-fetch) — ~15× faster AND kept verdict JSON out of orchestrator context. Validator (factual_verify.py status) enforces claim_id completeness so quality held (spot-checked batch 73 = real specific source URLs, not memory). STALE clusters into fact-packs (42/132 = 32% share a source_value): top = "350 users/anti-phishing policy" ×4 · OM3/OM4 fibre 300m/400m ×3 · Cloud Monitoring 375 projects ×3 · HA VPN no-SLA ×3 · _Default log retention 30d ×3 · gp3 80K IOPS ×2 · Snowball 210TB ×2 · DXGW 20-VGW ×2. STALE concentration by cert: gcp-network-engineer 13 · gcp-cloud-engineer 10 · juniper-jncis-ent 9 · aws-saa-c03 8 · gcp-security-engineer/pl-400 6. NEXT: triage + FIX the 132 STALE with Sush sign-off (Voice Rule) — fact-packs grep-replaceable; 90 singletons need per-Q edits; each fix touches live question JSON → Rule #1 SLA + test-guided-qa.cjs apply. Commits: c91058a (65-79,83-84) + 99ada73 (80-96 + strict merge). Helper: guided/files/_wrap_partials.py (injects manifest/prompt SHA + completed_at into agent-written partials).

Earlier — Last updated: 2026-06-03 NZT session d883c640 (4-phase back-to-back, Sush AFK 2h then briefly back to pick bug #28 path): 1. Bug #24 Cisco strip SHIPPED in e83a074 -- 90 Qs learnLink stripped (record undercounted 62 -> actual 90: CUCM cluster was 54 not 31). Per Rule #5 picked Option C (strip). Mashup scanner clean post-fix; SLA 3-check smoke 200. 2. P3 factual_check.py v0.1 SHIPPED in a487eb0 -- LAST of 4 P3 scanners. Rule #6 probe first; 952 unique HIGH Qs / 155 MEDIUM / 2595 LOW flagged. 10-hit calibration: ~10% scenario-fiction, ~80% verified-current, ~10% needs-LLM-verify. LLM verifier (factual_verify.py) NOT yet built -- 4-8h fresh-context work; handoff doc at session-state files. 3. ab-731 mid-tier triage DONE -- 20 unique HIGH Qs (record said 15). 0 mechanical fixes needed; d1-s07 RAG TCO logged as bug #28 with 4 Voice Rule rewrite paths for Sush (Path C recommended). 4. Bug #28 d1-s07 Voice Rule rewrite SHIPPED in 53e9920 -- Sush picked Path C inline. 6 fields rewritten (scenario/question/explanation/whyWrong[c]/whyWrong[d]/realWorld). Target swapped from $ amounts to input token volume (≥90% reduction; RAG hits 93%). Rubber-duck adopted 4 blocking findings. Bonus: factual_check.py sla_uptime regex hardened (caught self-spawned FP on '93% slashes').

Open bugs at session end: 3 (#1+#2 juniper deferred, #3 vault deferred — all three are deferred-by-Sush items; zero actionable Phase D bugs remain). Earlier same date session 97110a71: pricing sweep top-27 certs DONE -- 145 Qs triaged, 1 pricing bug #25 + 2 non-pricing bugs (#26 fixed, #27 expanded 1->385 Qs ALSO fixed). Cumulative pricing-TP rate 0.61% vs 40% diverse-sample baseline (66x off). Pricing sweep ESSENTIALLY COMPLETE. Update this doc at the end of every Phase D session (bump the date, move shipped items, tighten estimates).

Persistence model: bug capture lives in C:\ssClawy\guided\files\discovered_bugs.json (git-tracked, cross-session). Helper scripts in C:\ssClawy\guided\scripts\bugs\ (load_bugs.py + save_bugs.py). The per-session SQLite discovered_bugs table is loaded from / written back to this JSON at session boundaries. See "Discovery infrastructure" below.

TL;DR

You're continuing the Phase D portfolio mashup cleanup. 32 commits shipped so far (83cb705 + 3234ca0 + 91fe2cb + cc01005 + 2e75b81 + 3ebee72 + 49a668d + 80bc25c + d45e4d3 + 5992fa6 + d2b69e9 + af7fb22 + ea752df + 49f0cf6 + fa885af + 66934b5 + bcefb0a + 71f3668 + fcde55c + 3543d17 + c06cc4a + 21cc22e + 4b739d6 + c7f7bde + 42f4e37 + 50e4768 + d934cdc + e873236 + cadb7b0 + baf335c + 85e97ee + playbook-update-pending). P1+P7+P8+P9+P10 DONE. P2 essentially complete (~98%). P3: 3 of 4 scanners DONE; only factual_check.py LLM-based remains. #25 ai-200 SHIPPED this session (97110a71) — Premium ACR explanation invented ~$1.50/GB/month rate structure that doesn't exist (real: $1.667/day flat + 500 GiB included + $1.667/day per geo-replicated region). Voice Rule answer (a) STRIP. 17 Qs triaged across top-5 GSC certs (ai-200/dp-750/ai-300/sc-500/az-140); 1 real bug + 14 SCENARIO_FICTION + 2 DRIFTABLE_OK = 5.9% real-bug rate on top-GSC vs 40% diverse-sample baseline. New heuristic: top-GSC certs are newer Microsoft builds with less drift. 🆕 Open bugs as of session end: 4 (unchanged — #1+#2 juniper deferred, #3 vault deferred, #24 Cisco residual LOW; bug #25 added AND closed in same session).


Current state

What's shipped (live on prod, SLA green)

Commit Author session What
83cb705 9e5cc23f 5 H1 mashup fixes (eccouncil-cnd-v3 ×4 + hashicorp-vault ×1)
3234ca0 9e5cc23f SME-flagged factual fix in vault d4-008 (correct=b → c per Vault docs)
91fe2cb 89327d9c B-Lite strip on 45 Qs + 1 deletion (juniper-jncis-ent + 4 cross-cert) + new portfolio_mashup_scan_v2.py with H4 heuristic
cc01005 884a05e7 P1 DONE — wired mashup scanner into inject_phase_b.py as pre-commit gate. H1+H4 hard fail; H2+H3 warn. --strict-mashup elevates H2. --check-only for CI. Tested with 4 synthetic payloads.
2e75b81 7a664e44 3 H4-only HIGH fixes (gcp-pca-d3-033, gcp-pde-d1-026, juniper-jncia-junos-d3-023) + new template_boilerplate_check.py (probe-only). 5 H4-only HIGH confirmed as FPs (62.5% FP rate cross-cert).
3ebee72 7a664e44 P7 portfolio strip DONE — 1572 template boilerplate fragments stripped from 1221 Qs across 57 certs (~4% of 31,212-Q bank). New strip_template_boilerplate.py script. Mashup scanner re-run: 0 new mashup hits introduced; cleared 2 H4 HIGH + 36 H4m MEDIUM incidentally.
49a668d 7a664e44 P8 template-boilerplate prevention hook DONE (B1-B6) — wired template_boilerplate_check.scan_question into inject_phase_b.py alongside the mashup hook. All B1-B6 hits HARD-FAIL. Tested with 3 synthetic payloads (clean / B1+B2+B3 / B4+B5+B6).
80bc25c c0952d49 P8.5 scanner+stripper extension — added B7-B12 phrase entries to template_boilerplate_check.py + strip_template_boilerplate.py. New strip_question_stem() code path for qstem_only entries (end-anchored regex prevents FPs on honest mid-sentence uses). Help text in inject_phase_b.py updated to list B1-B12.
d45e4d3 c0952d49 P8.5 content strip on az-104 — 288 mechanical text removals across 48 Qs in az-104-domain-3.json: 268 field-level (B7-B11) + 20 question-stem (B12). All pure text deletion; no semantic content modified. Mashup scanner delta: HIGH 328→328, MEDIUM 6466→6464 (2 incidental cleanups + 1 ratio-drift artifact on d3-045 H4m).
5992fa6 c0952d49 data: discovered_bugs.json bugs 9+10 (az-104 B7-B11 cluster + B12 qstem) marked fixed.
d2b69e9 c0952d49 mb-800 d4-034 + d4-035 deletions — both corrupt-options mashup (Q stem swapped, options/correct/explanation/whyWrong tell different stories). Same class as juniper-jncis-ent-d2-008. Per Phase D rule (under-representation > misleading content): deleted. d4 Qs 67 → 65.
af7fb22 c0952d49 data: discovered_bugs.json bugs 11+12 (mb-800 deletions) marked fixed.
ea752df 1ef77611 P9 marker-swap scanner + prevention hook DONE — new marker_swap_check.py detects whyWrong[X] (X NOT in correct) containing verbatim CAPS "This IS the correct" / "This IS the right" — author-artifact signal that correct marker was changed without updating whyWrong. Calibrated 100% real-bug rate (4/4 portfolio hits real). Wired into inject_phase_b.py as HARD-FAIL gate.
49f0cf6 1ef77611 pl-300 marker-swap + enrichment-contamination fixes — 4 Power BI authoring Qs (d1-012, d1-013, d1-016, d1-021) had correct marker pointing to wrong option while explanation + whyWrong supported the actual right answer. All 4 ALSO had hint/examTip/realWorld copy-pasted from totally unrelated Qs. Fix per Q: swap correct marker (no content invented — the right answer was already in the options), rewrite whyWrong for the now-wrong option, B-Lite strip off-topic enrichment. Scanner delta: pl-300 HIGH 6→2 (-4), portfolio HIGH 326→322 (-4).
fa885af 1ef77611 data: discovered_bugs.json bugs 13-16 (pl-300 marker-swap class) marked fixed.
66934b5 3352b87d ab-731 enrichment-contamination-without-marker-swap fix — d2-053 (Copilot Studio + Power Automate). Marker was correct ([b]), but whyWrong[a]/[c]/[d] + examTip were copy-pasted from a different Q discussing Power BI / Azure DevOps / SharePoint workflows (distractors that aren't even in this Q's options). NEW SUBCLASS distinct from pl-300 marker-swap (no marker swap needed; just contaminated enrichment). Escapes H4 because vocabulary overlap (workflow / automation / integration) keeps H4 ratio above 0.07. Rewrite whyWrong + examTip from clean source (option text + explanation, same Voice precedent as pl-300 rewrites).
bcefb0a 3352b87d gcp-pcdb-d2-048 deletion — merged mashup of "DR compliance" Q (scenario + Q stem + explanation + hint + examTip + realWorld + learnLink) + "Table bloat" Q (options + correct=[a] + whyWrong[b/c/d]). Right answer for "DR compliance" not in options; right answer for "table bloat" not asked. Same class as juniper-jncis-ent-d2-008 + mb-800-d4-034/035. DELETED per corrupt_options rule. gcp-database-engineer-d2 67→60 Qs (was 61 in scan; off-by-one in scan output but verified via raw file inspection — 61 was correct pre-delete, 60 post-delete).
71f3668 3352b87d data: discovered_bugs.json bugs 17-18 (this session's finds) + bugs 19-20 (P10/P11 tooling candidates) + hygiene fix on bugs 13-16 (fix_commit "pending" → "49f0cf6").
fcde55c 7bb2af7a P10 shipped — H2 reclassified HIGH→MEDIUM in portfolio_mashup_scan_v2.py (single-line surgical change). Portfolio HIGH dropped 320→6 exactly as calibration predicted (97.5% FP rate over 320 hits / 97 certs).
3543d17 7bb2af7a data: discovered_bugs.json bug #19 marked fixed (P10 = fcde55c).
c06cc4a 7bb2af7a P11 WONTFIX with probe evidence — built p11_probe.py calibration probe (5 known-real drift + 50 H2FP + 200 HONEST × 5 token-overlap metrics). No natural cliff exists; REAL drift cases sit inside HONEST distribution on every metric. ab731-d2-053 subclass now covered by P3-time manual SME audit. Probe data preserved at files/p11_probe_data.json so future sessions don't re-invent.
21cc22e 7bb2af7a P12 NEWlearnlink_validator.py (async HEAD-check across 7,153 unique URLs in ~5 min). 1,293 dead URLs (18%) affect 2,691 Qs (~8.6% of bank). Top concentration: 200 Qs hit ONE dead CompTIA URL. Bug #21 logged with 3 candidate remediation paths.
4b739d6 003b7095 P3 scanners (tooling)pricing_check.py + deprecated_check.py shipped + Rule #6 probe templates (files/pricing_probe.py + files/deprecated_probe.py) + bug-21-decision-support helper (files/learnlink_priority_rank.py) + bug-22 reproduction (files/find_legacy_gpt_pricing.py) + triage helper (files/pricing_triage_sample.py). aws_sdk regex calibrated mid-session (FP fix: 13 hits eliminated).
c7f7bde 003b7095 P3 data + bugs — baseline pricing/deprecated scans + dead-link priority ranking (Tier 1=6 certs/115 Qs, Tier 4=48 certs/1,769 Qs); bug #21 enriched with prioritization context; bug #22 (stale GPT-4 token pricing cluster in ab-731-d1, 5 Qs); bug #23 (terminology aging summary, 0 HIGH / 225 MEDIUM portfolio-wide).
42f4e37 003b7095-step1 #23 Entra rename SHIPPED (Path B selective) — 14 files / 13 Qs / 37 renames. 2-pass audit caught 15 Qs with intentional historical use (formerly Azure AD, legacy Azure AD, Microsoft Entra ID (Azure AD) parenthetical) and reverted those. Bug #23 closed.
50e4768 003b7095-step2 #21 Tier 1+3 dead-link sweep SHIPPED (Path A') — 26 files / 441 learnLink edits. 8 manual overrides (MS Copilot Studio multi-agent reorg, CompTIA Data+ /dataai->/data, 4 NIST /pubs/->/publications/detail/, Cisco CCNA /s/article->/s/ccna, Redis distributed-locks slug). 36 auto-chain replacements with HUB_ALLOWLIST + TOO_GENERIC blocklist preventing near-domain-root replacements. 3 Cisco URLs (62 Qs) couldn't find safe replacements → logged as bug #24. Bug #21 closed.
d934cdc 003b7095-step2 data(phase-d): bug record updates (#21+#23 fixed; #22 enriched with Build 2026 context for voice-prep next session; #24 NEW for Cisco residuals).
e873236 b5cb9fa4 #22 ab-731 SHIPPED (Path C license-vs-credits rewrite) — 5 Qs in ab-731-domain-1 (d1-019, d1-020, d1-021, d1-022, d1-s06) pivoted from stale March-2023 GPT-4 token pricing ($0.03/$0.002 per 1K) to Build 2026 M365 Copilot architecture. Voice Rule answer (c) — decision tree per use-case. Rubber-duck-validated: E7 removed from d1-s06 numeric math (no sourced price); Credits never positioned as substitute for human-Copilot licence; workflow-based decisions not role-label-based; "buckets are additive" gotcha in d1-022; new "Credits-only" wrong answer in d1-s06 testing the key misconception; brand-voice scrub. Zero new mashup HIGH; portfolio MEDIUM dropped 6781→6780 incidentally. All 5 learnLinks validated 200.
cadb7b0 b5cb9fa4 data(phase-d): bug #22 closed (with cert-identity correction note — ab-731 is 'AI Transformation Leader' not 'Copilot Sales Specialist' as the bug record originally claimed).
baf335c 97110a71 #25 ai-200 SHIPPED (Voice Rule answer (a) STRIP) — ai200-d1-s01 Premium ACR explanation invented ~$1.50/GB/month × 3 regions × ~80 GB rate structure that doesn't exist (real: $1.667/day flat + 500 GiB included + $1.667/day per geo-replicated region). Deleted the misleading sentence; teaching point on 3-features-on-ONE-registry preserved. JSON valid; pricing scanner re-run ai-200 HIGH dropped 2 Qs → 1 Q; mashup scanner HIGH unchanged 6→6; learnLink 302→200; production cache-busted verified.
85e97ee 97110a71 data(phase-d): bug #25 closure + bug #22 fix_commit leftover fix (PENDING → e873236). 32 commits cumulative; 5 bugs closed; 4 open (unchanged). Pricing sweep follow-on STARTED: 17 Qs triaged across top-5 GSC certs (ai-200/dp-750/ai-300/sc-500/az-140); 5.9% real-bug rate (vs 40% diverse-sample baseline) — top-GSC certs are newer + cleaner.

Scope numbers (refined)

  • 126 certs · 31,208 questions · 247 Qs/cert avg
  • Scanner-flagged HIGH (current, post-P10): 6 (mashup) + 0 (deprecated) + 325 unique Qs (pricing, mostly verifiable claims, 40% TP rate on triage sample)
  • Scanner-flagged MEDIUM (current, post-P10): 6,781 mashup + 225 deprecated (mostly Azure AD rename) + 529 unique Qs pricing (bare-dollar scenario amounts)
  • Template boilerplate (B1-B12): 0 remaining (1944 fragments cleared across 3ebee72 + d45e4d3)
  • Marker-swap (CAPS "This IS"): 0 remaining portfolio-wide (4 fixed in 49f0cf6)
  • Dead learnLinks: 1,293 unique URLs affecting 2,691 Qs (~8.6% of bank). Prioritization: Tier 1 (GSC top-10) = 6 certs/115 Qs ← actionable customer-impact set. Tier 4 (zero-GSC-traffic) = 48 certs/1,769 Qs = 66% of total, safe to defer. See bug 21 (open).
  • ~~🆕 Stale vendor pricing: 5 Qs in ab-731-d1 with legacy GPT-4 $0.03/1K token pricing references (bug 22 open). Concentrated; no spread to other certs.~~ FIXED in session b5cb9fa4 — 5 Qs rewritten with Build 2026 license-vs-credits framing.
  • 🆕 Deprecated patterns: 0 HIGH portfolio-wide (clean!) + 225 MEDIUM aging-terminology hits (Azure AD rename cluster, gsutil refs). See bug 23 (open).
  • P2 essentially complete — ~98% portfolio coverage. 97.5% portfolio-wide FP rate confirmed.
  • H4 cross-cert FP rate: 62.5% (5/8 in 7a664e44 batch) — high-precision relative to H1/H2 but still mostly cross-cert FPs

Triage progress (P2 — 320/316 H1/H2 HIGH = ~100% portfolio-wide, 97 certs done. Remaining 29 certs have ZERO H1/H2 HIGH and are pristine.)

Per-cert triage table archived (was getting long). Summary:

Tier Certs triaged H1/H2 HIGH Real bugs
24-cert first wave (sessions 9e5cc23f → 1ef77611) 24 165 6 (2 mb-800 deletions + 4 pl-300 marker-swap fixes)
4-HIGH tier (this session 3352b87d) 7 28 0
3-HIGH tier (this session 3352b87d) 15 45 0
2-HIGH tier (this session 3352b87d) 25 50 2 (ab731-d2-053 enrichment rewrite + gcp-pcdb-d2-048 deletion)
1-HIGH long tail (this session 3352b87d) 32 32 0
TOTAL 97 320 8 real + 2 cross-cert template strips from earlier sessions = 10 portfolio-wide

Per-cert detail (which 4-/3-/2-/1-HIGH certs were triaged this session) in ~/.copilot/session-state/3352b87d-144b-4d45-8149-1087b14853f9/files/triage-*.txt if forensic detail is ever needed. Headline: every batch this session was overwhelmingly H2 contrast-to-right-answer FPs.

What's open (priorities in order)

🎯 Quality bar (set 2026-06-03 by Sush, session d883c640): Every paid cert gets fixed, regardless of sales or GSC traffic. Phase D defers items based on prioritisation against other higher-value work, never based on "this cert doesn't sell." If a Q is on the platform and someone could buy that cert at `$9, it deserves the same quality bar as az-900 / sc-500 / ai-200. Atlas must not characterise low-traffic certs as "deferred safely" or "dead cert" — that's a misread of the product standard.

🎯 Priority discipline (locked 2026-06-03 session b5cb9fa4 with Sush sign-off): drain the well before drilling the next one. When a discovery scanner ships, its full TP harvest comes BEFORE building the next scanner. Bug #22 was the most concentrated cluster of pricing_check.py's 325 HIGH hits / ~130 estimated real bugs portfolio-wide; the remaining ~125 customer-facing pricing bugs sit in higher-traffic certs and outrank both the LOW-impact #24 cleanup AND building factual_check.py from scratch. Today's d1-s07 finding (RAG TCO math with a 0.03 reference NOT caught by the bug-22 canonical detector) is direct evidence more bug-22-equivalents are waiting in the HIGH pool.

  1. 🆕 Pricing sweep follow-on (P3.1 TP harvest, ESSENTIALLY COMPLETE as of session 97110a71) — Top-27 GSC certs DONE this session (145 Qs triaged across 27 certs spanning every major vendor — Microsoft AI/Azure/M365/Power Platform/Dynamics, AWS Cloud/Dev/Architect/ML, GCP Digital Leader, Cisco, Fortinet, CompTIA, ISC2, ISACA, EC-Council). 1 real pricing bug shipped as #25 (ai-200 Premium ACR) + 2 non-pricing bugs found (#26 aws-mla-c01-d4-036 multi-issue contamination FIXED in 54978f7+003bcc4; #27 gcp-cdl-s6-008 cross-cert template contamination OPEN). 142 FPs (77 SCENARIO_FICTION + 51 DRIFTABLE_OK with all vendor-price claims verified current). Cumulative pricing-TP rate 0.69% (1/145) vs 40% diverse-sample baseline — calibration was 58x off. aws-clf-c02 stress test (40 HIGH Qs — the pricing-heaviest cert in the portfolio where pricing IS the core domain) returned ZERO real bugs. See "Pricing-triage register" below for the full 27-cert table. What's left: ~155 HIGH Qs sit in long-tail certs with zero recent GSC traffic (older AWS-PAS/SCS/MLS/ANS + GCP-DB-eng/Architect/Dev + many vendor-specific certs). Given the 0.69% rate, expected real-bug yield is 1-2 bugs across 155 Qs = ~13 hours of triage for ~1 bug = NEGATIVE ROI. Recommend: DECLARE PRICING SWEEP COMPLETE and move on to (a) bug #27 quick fix + portfolio-wide 'AWS in non-AWS hint/examTip' grep cleanup, (b) #24 Cisco strip, or (c) P3 factual_check.py last scanner build. The 'older AWS/GCP cluster has all the drift' hypothesis from earlier sessions is empirically DISPROVEN — aws-clf-c02 was the perfect test case and it was clean. Tooling: GSC × pricing intersection script lives at ~/.copilot/session-state/97110a71-0b79-4769-80b2-1e5bb4a90716/files/gsc_guided_pricing_intersect.py (re-runnable for fresh GSC pulls). Triage rubric calibrated: STALE_REAL_PRICE (fix per Voice Rule) / SCENARIO_FICTION (narrative customer-story math — FP, no action) / DRIFTABLE_REAL_CITATION (verify against vendor docs, log as VERIFIED_OK in this playbook's register so future sessions skip).
  2. ~~#24 NEW: 3 Cisco deep-doc URLs (62 Qs in cisco-clcor/ccie-ei/spcor) — auto-chain couldn't find safe replacements. LOW severity (all in zero-GSC-traffic certs). Three options: manual research per URL, replace with search-query URL, or strip learnLink. Recommend strip. Good buffer task between heavier pricing-sweep sessions (30-60 min slot).~~ ✅ SHIPPED e83a074 (session d883c640). Picked Option C (strip). Re-count: 90 Qs actually affected (bug record undercounted by 28; CUCM cluster was 54 not 31, VXLAN was 20 not 15). Helper script files/bug24_strip_cisco_links.py (idempotent — re-runnable for any future Cisco dead-link cycle).
  3. P3: Build last remaining new discovery scanner — ~~factual_check.py (LLM-based, ~4-8h). Other 3 P3 scanners (learnlink_validator + pricing_check + deprecated_check) DONE. Defer until pricing-sweep well is drained to avoid accumulating bug debt while discovering new bug categories.~~ ✅ Scanner v0.1 SHIPPED in a487eb0 (session d883c640). Pattern-only regex scanner, no LLM yet -- 952 unique HIGH Qs / 155 MEDIUM / 2595 LOW flagged portfolio-wide. 10-hit calibration: ~10% scenario-fiction, ~80% verified-current, ~10% needs-LLM-verify. NEXT: build factual_verify.py LLM verifier in fresh session (4-8h) — full handoff with architecture options + verifier prompt template at ~/.copilot/session-state/d883c640-6bd7-40da-85e4-2bd18f9abed0/files/p3-factual-handoff.md. Recommended path: research task agent (no per-call cost, Microsoft Learn + web grounding).
  4. ~~#28 NEW (DEFERRED — Voice Rule required) — ab731-d1-s07 RAG TCO Q uses OLD GPT-4 token pricing ($0.03/$0.06 per 1K) to compute $28,500/day baseline. Math internally consistent but 2026 GPT-4o pricing is ~6× cheaper for input / ~3× cheaper for output (real customer would see ~$5,500/day baseline). Architectural lesson (RAG vs full-context) is timeless; numbers are 2023 window-dressing. Sister cluster to bug #22. 4 rewrite paths surfaced in bug #28 record; recommended Path C (abstract ratios). This was the latent finding from b5cb9fa4 #22 session.~~ ✅ SHIPPED 53e9920 (session d883c640 follow-on, Sush picked Path C inline). 6 fields rewritten (scenario, question, explanation, whyWrong[c], whyWrong[d], realWorld). Target swapped from dollar amounts to input token volume (90% reduction target; RAG hits 93%). Rubber-duck-validated with 4 blocking adoptions. Bonus: factual_check.py sla_uptime regex hardened with word boundary (caught my own scanner FP on '93% slashes' in the new explanation). Open bugs 4 → 3.
  5. P4: Universal per-cert SME audit as standard step during P3 triage — catches non-mashup bugs that no heuristic can detect. Now also covers the ab731-d2-053 subclass (enrichment_contamination_without_marker_swap) that P11 could not detect heuristically.
  6. P5: Juniper re-author 41 stripped Qs — Bug #2 MEDIUM. 41 Qs in juniper-jncis-ent had enrichment blocks (whyWrong/hint/examTip) stripped in commit 91fe2cb during the original juniper B-Lite enrichment cleanup. Qs are FUNCTIONAL (stem + options + correct + explanation intact); just sparse enrichment vs other certs. Voice Rule rewrite — ~4-8h. Not safely skippable per the 'every paid cert gets fixed' standard — schedule for a dedicated session like we did for ab-731 #22.
  7. P6 (defer): MEDIUM triage timeboxed on top-20 GSC. Likely ~95% FP.

🆕 Phase E (planned follow-on after Phase D fully closes): enrichment audit + restoration across all 126 certs

Trigger: Sush set 2026-06-03 session d883c640 — "we have to enrich all certs after this work we are doing."

Scope candidates (needs Sush scope-pick at Phase E kickoff): - Narrow: Restore the 41 juniper Qs to full enrichment quality (= bug #2 above) - Medium: Audit + improve enrichment across top-20 GSC certs — ensure every Q has whyWrong / hint / examTip / realWorld of consistent quality - Wide: Same audit across all 126 certs — estimated 8-15 sessions at 30 min/cert

SOP candidates: - Build an enrichment_check.py scanner that flags Qs missing any of {whyWrong, hint, examTip, realWorld} or with placeholder/templated content - Per-cert SME audit (re-use Phase D pattern: research agent + voice-aligned rewrite) - Voice Rule sign-off per cert (large customer-facing surface)

Not Phase D scope. Phase E kickoff = after Phase D fully closes (factual_verify.py shipped + #1 + #2 + #3 all fixed). Sush picks the scope at kickoff.

Latent finding for the pricing sweep: ab731-d1-s07 (RAG architecture cost-reduction Q) has a "0.03" reference NOT caught by find_legacy_gpt_pricing.py (different pattern: TCO math, not GPT-token cluster). ~~Needs hand-triage during the next ab-731 pricing-sweep session.~~ ~~NOT triaged in this sweep — ab-731 mid-tier batch deferred (still has 15 HIGH Qs including the #22-adjacent d1-s07; not in top-27 GSC slice this session because ab-731 sits at "1 click + 1559 impr" with 5 of its 15 HIGH Qs already fixed by #22). Worth a focused next-session pass since ab-731 IS where the only confirmed real bug pattern lives (#22 was the only customer-facing rewrite needed).~~ ✅ TRIAGED in session d883c640 (20 unique HIGH Qs, not 15 — bug record undercount). d1-s07 confirmed STALE_REAL, logged as bug #28 with 4 Voice Rule rewrite paths for Sush. 0 mechanical fixes needed (other 19 Qs are VERIFIED_OK Copilot $30/u/mo current refs, scenario-fiction, or defensible Azure scenarios).

Triage learning from session 97110a71 (CONFIRMED across 145 Qs / 27 certs in one session, including the pricing-heaviest stress-test cert): estimated real-bug rate of 40% (from initial 25-Q diverse-sample calibration) is dramatically wrong for any traffic-weighted slice — actual rate was 0.69% (1/145) spanning every major vendor and the AWS Cloud Practitioner stress test (40 HIGH Qs, 0 real bugs). The 40% calibration was a sampling artefact, not a true portfolio rate. The "older AWS/GCP has the drift" hypothesis is DISPROVEN — aws-clf-c02 (Cloud Practitioner) was the perfect test case for that hypothesis (pricing IS the core domain) and it was 100% clean. Final revised estimate: ~5-10 real bugs across the entire 31,212-Q portfolio, mostly clustered in cert-specific scenario types (e.g., ab-731's GPT-token decision Qs were the #22 cluster; ai-200's Premium ACR was #25; no systematic vendor-pricing drift detected anywhere else). Bonus findings: 2 non-pricing bugs surfaced (#26 + #27) by reading Qs during pricing triage — validates the P4 SOP step "per-cert manual SME audit during P3 triage". Strategic recommendation: declare pricing sweep complete and pivot to P3 factual_check.py (which may find more bugs than this remaining pricing-sweep tail), or to bug #27 + portfolio-wide template-contamination grep cleanup.

Pricing-triage register (session 97110a71)

These 48 Qs have been manually verified and don't need re-triage on future pricing_check.py runs unless the underlying Microsoft/AWS pricing pages change materially. Future sessions can skip these and focus on un-triaged certs.

Cert Triaged Qs Real bug FPs (sceanrio-fiction / verified-current) Verdict
ai-200 d1-s01, d2-014 1 (d1-s01, FIXED in baf335c) 0 / 1 (OpenAI embedding $0.02/$0.13/1M) DONE
dp-750 d4-029, d4-059 0 2 / 0 DONE
ai-300 d2-021, d2-032, d4-023, d5-s01 0 4 / 0 DONE
sc-500 d2-023, d3-s03, d4-010, d4-044, d4-s03 0 2 / 3 (Sec Copilot SCU, Defender for Servers $15/VM/mo Plan 2, scenario-fiction Sentinel) DONE
az-140 d1-036, d1-042, d1-044, d1-078 0 4 / 0 DONE
az-104 d3-009, d3-014, d3-028, d4-s08, d4-003, d4-037 0 1 / 5 (D4s_v5 $140/mo, Std SSD $38 vs Premium $73, PE $0.01/hr math, VPN Gw $140+/mo) DONE
comptia-cas-005 d1-001, d1-007, d2-042 0 3 / 0 DONE
comptia-cy0-001 d2-s11, d3-s06 0 2 / 0 DONE
isc2-cissp-issep d2-013 0 1 / 0 DONE
aws-dva-c02 d2-021, d2-039, d2-052, d2-s09, d4-s01, d4-s12 0 0 / 6 (KMS $1/mo, SecretsMgr $0.40/secret/mo, Param Store Advanced $0.05/mo, ProvConc GB-s, PutMetricData $0.30/M+$0.30/metric/mo) DONE
aws-sap-c02 d2-004, d2-s10, d3-014, d3-018, d3-019, d3-020, d3-021, d3-046 0 6 / 2 (CloudFront Funcs $0.10/M, Shield Adv $3,000/mo) DONE
az-400 d3-028, d3-099 0 2 / 0 DONE
dp-800 d3-011 0 1 / 0 DONE
isaca-cism d2-018, d4-010 0 2 / 0 DONE
az-900 d1-009, d1-010, d1-025, d1-s02, d1-s07, d1-s09, d1-s10, d2-015, d2-025, d2-066, d2-075, d2-s03, d2-s04, d2-s09, d2-s12, d2-s14, d2-s17, d3-001, d3-s09, d3-s13 0 16 / 4 (Private Endpoint $7/mo, VPN Gw1 $130/mo, ExpressRoute $2k+/mo, App Service F1 $0 + PostgreSQL B1ms $12-15/mo) DONE (2 borderline narrative numbers flagged: d2-s12 M416s at $2.8k/mo and d3-s09 $2.8k compute jump are stretched but defensible scenario math)
aws-mla-c01 d1-s13, d2-008, d2-033, d2-038, d2-047, d2-050, d3-006, d3-035, d3-036, d3-042, d3-s01, d3-s12, d4-020, d4-036, d4-s12 0 (pricing) 3 / 11 (ml.p3.2xlarge $3.83/hr, ml.p3.8xlarge $14.69/hr, ml.g4dn.xlarge $0.736/hr, ml.m5.xlarge $0.23/hr EC2 + $0.269/hr SageMaker hosting, ml.m5.large $0.134/hr, Spot 60-90%, Inferentia 50-70% savings) DONE — BUG #26 LOGGED + FIXED in 54978f7 + 003bcc4 (d4-036 multi-issue contamination repair: type mcq, correct=['a'], stripped option (e) garbage, added whyWrong[b]+[c], fixed whyWrong[d] warehouse-mgmt phantom reference)
pl-900 d1-s07 0 0 / 1 (Power Automate Premium $15/u/mo + Power Apps Premium $20/u/mo) DONE
pl-400 d1-s09 0 1 / 0 DONE
isc2-ccsp d6-026 0 1 / 0 DONE
fortinet-nse4 d4-026 0 1 / 0 DONE
cisco-clcor d3-026 0 1 / 0 DONE
comptia-cv0-004 d3-s06, d4-s06 0 2 / 0 DONE
comptia-sy0-701 d5-003, d5-005, d5-027 0 3 / 0 (all quantitative-risk ALE pedagogy) DONE
isc2-cissp-issmp d3-024, d3-028, d5-009, d5-016 0 4 / 0 (all FAIR/quantitative-risk + DR pedagogy) DONE
gcp-cloud-digital-leader s1-003, s1-010, s1-022, s6-008 0 (pricing) 4 / 0 DONE — BONUS BUG #27 LOGGED: s6-008 hint mentions "which AWS service" in a GCP cert (cross-cert template contamination, LOW severity, OPEN)
aws-clf-c02 d1-004, d1-022, d1-035, d1-036, d1-040, d1-s04, d1-s08, d1-s13, d2-014, d2-028, d2-039, d3-019, d3-020, d3-036, d3-039, d3-s04, d3-s05, d3-s09, d4-001, d4-003, d4-004, d4-006, d4-007, d4-008, d4-009, d4-011, d4-016, d4-019, d4-020, d4-022, d4-024, d4-s01, d4-s02, d4-s03, d4-s05, d4-s06, d4-s07, d4-s08, d4-s09, d4-s11 0 18 / 8 (t3.medium $0.0416/hr, Athena $5/TB scanned, cross-AZ $0.01-0.02/GB, cross-Region $0.02-0.09/GB, Business Support $100/mo min, Enterprise Support 15-min critical response + tier math, EC2 egress Tokyo $0.114/GB, S3 tiered storage $0.023/0.022/0.021/GB at 50/450/500+ TB) DONE — STRESS-TEST PASSED. Pricing-heaviest cert in portfolio (40 HIGH Qs, AWS Cloud Practitioner where pricing IS the core domain). Zero real bugs. Decisively disproves the 40% diverse-sample calibration.
ab-731 (mid-tier batch, session d883c640) d1-019, d1-020, d1-021, d1-022, d1-024, d1-033, d1-s06, d1-s07, d2-020, d2-024, d2-025, d2-028, d2-066, d2-s01, d2-s03, d2-s06, d2-s16, d3-042, d3-045, d3-049 1 (d1-s07 STALE_REAL, DEFERRED as bug #28 for Sush Voice Rule rewrite) 4 / 15 (5 of 20 already fixed by bug #22 in e873236: d1-019/020/021/022/s06 carry current $30/u/mo refs; 8 more Qs with $30/u/mo Copilot price are VERIFIED CURRENT 2026; d3-042/045/049 with $30+committed-use refs DRIFTABLE_OK; d1-024/033/d2-020/066 are scenario-fiction; d2-s06 $5K/mo Azure budget + d3-049 $5-10K/mo CUD threshold defensible) DONE — d1-s07 is the only customer-facing rewrite needed (latent finding from bug #22 b5cb9fa4 session). 4 Voice Rule rewrite paths surfaced in bug #28 record. ab-731 mid-tier was the "only remaining traffic-meaningful pricing cluster" per Sush.
TOTAL 165 2 pricing (1 fixed #25 + 1 deferred Voice Rule #28) + 2 non-pricing (#26 fixed, #27 expanded to 385-Q portfolio strip ALSO FIXED in 1c07f76+ad029bc) 81 / 66 0.61% pricing TP (1 fixed of 165)

🔥 BIG WIN — Bug #27 expanded scope discovery + fix (post-sweep, same session): What started as a single-Q gcp-cdl-s6-008 contamination report expanded — via portfolio-wide grep — into a SYSTEMIC 385-Q content bug across 6 different non-AWS cert families (az-305 165 Qs, gcp-cloud-digital-leader 62, gcp-security-engineer 58, dp-300 55, isc2-cc 32, sc-300 13). Identical hint template "Consider the primary access pattern and constraints described in the scenario — which AWS service is purpose-built for that specific pattern?" had leaked into all 385 Qs during original authoring. Customers reading Azure Architect (az-305) Qs were being told to think "AWS service" when the correct answer is Azure. Pure mechanical strip applied (no content invention) — 15 files modified, 430 deletions vs 45 insertions (385 hint strips + JSON reformat). 8 hits intentionally NOT stripped (legitimate cross-cloud questions in Cisco SCOR multi-cloud security, CompTIA CASP+ forensics, CEH cloud pentesting, CHFI forensics — these certs DO cover AWS as part of their scope). Severity upgraded LOW → HIGH on scope discovery. Production verified post-CF-deploy (6 min lag for the 15-file build). This is exactly the kind of portfolio-scale content bug the pricing sweep inadvertently surfaced via P4 (per-cert manual reading) — validates P4 as a mandatory SOP step. The 145-Q pricing triage found 1 real pricing bug but 2 non-pricing bugs (#26 + #27), and #27 alone was 385x bigger than the original 1-Q report.

What's DONE (shipped this week)

  • ✅ ~~P1: Mashup prevention hook in inject_phase_b.py~~ — commit cc01005 (884a05e7 session)
  • ✅ ~~Scanner refinement v2 (H4 heuristic)~~ — commit 91fe2cb (89327d9c session)
  • ✅ ~~Juniper B-Lite enrichment strip (45 Qs)~~ — commit 91fe2cb (89327d9c session)
  • ✅ ~~Top-20 GSC HIGH triage (78 findings)~~ — commit 83cb705 + decisions (9e5cc23f session)
  • ✅ ~~Vault d4-008 SME-flagged factual fix~~ — commit 3234ca0 (9e5cc23f session)
  • ✅ ~~H4-only HIGH triage (8 findings)~~ — commit 2e75b81 (7a664e44 session)
  • ✅ ~~P7: Portfolio strip of Phase B template boilerplate B1-B6 (1572 fragments)~~ — commit 3ebee72
  • ✅ ~~P8: Template-boilerplate prevention hook B1-B6~~ — commit 49a668d
  • ✅ ~~P8.5: Scanner+stripper extension to B7-B12 + az-104 secondary strip (288 fragments)~~ — commits 80bc25c + d45e4d3 (c0952d49 session)
  • ✅ ~~P2 round 1: First 6 certs triaged (63 HIGH, 96.8% FP rate, 2 real bugs deleted in mb-800)~~ — commits d2b69e9 (c0952d49 session)
  • ✅ ~~P9: Marker-swap scanner + prevention hook + 4 pl-300 fixes~~ — commits ea752df + 49f0cf6 + fa885af (1ef77611 session)
  • ✅ ~~P2 round 2: 12 more certs triaged (102 HIGH, 96.1% FP rate, 4 marker-swap real bugs fixed in pl-300)~~ — commits ea752df + 49f0cf6 + fa885af (1ef77611 session)
  • ✅ ~~P2 mass-triage (essentially complete): 79 more certs across 4/3/2/1-HIGH tiers + long tail (155 H1/H2 HIGH, 98.7% FP rate this batch; portfolio cumulative 320/316 H1/H2 HIGH triaged = ~100%, 97.5% portfolio-wide FP). 2 real bugs fixed: ab731-d2-053 enrichment-contamination-without-marker-swap rewrite + gcp-pcdb-d2-048 corrupt-options deletion.~~ — commits 66934b5 + bcefb0a + 71f3668 (3352b87d session)
  • ✅ ~~P10: H2 mashup hits reclassified HIGH→MEDIUM (320→6 portfolio HIGH; calibrated 97.5% FP rate over 320 hits / 97 certs). Frees triage cycles for P3.~~ — commit fcde55c + 3543d17 (7bb2af7a session)
  • ✅ ~~P11: Marked WONTFIX with probe evidence (5 known-real drift + 50 H2FP + 200 HONEST samples × 5 metrics tested). Token-overlap cannot discriminate drift from honest divergence. ab731-d2-053 subclass now covered by P3-time manual SME audit.~~ — commit c06cc4a (7bb2af7a session)
  • ✅ ~~P12 tooling (learnlink_validator.py): async portfolio HEAD-checker — 7,153 unique URLs in ~5 min, classifies dead/redirect_minor/redirect_major/server_error/network_error/client_error/ok. Run portfolio-wide; surfaced 1,293 dead URLs / 2,691 Qs affected as bug 21 for strategic call.~~ — commit 21cc22e (7bb2af7a session)
  • ✅ ~~P12.5 prioritization (this session): dead-link × GSC cross-rank. Tier 1 (GSC top-10) = 6 certs/115 Qs (actionable); Tier 4 (zero-traffic) = 1,769 Qs/48 certs (defer). Path A' = Tier 1 + Tier 3 top-10 = ~500 Qs for ~30 URL fixes recommended sub-path. Locale-strip hypothesis tested and dead (0/38 /en-us/ URLs returned 200 after strip).~~ — commit c7f7bde (003b7095 session)
  • ✅ ~~P3.1 (this session): pricing_check.py scanner — HIGH (dollar+rate-context) / MEDIUM (bare-dollar) / LOW (free-tier, savings-plan). 797 HIGH hits / 325 unique Qs portfolio-wide. 40% TP rate on 25-Q diverse triage sample (60% are scenario-fiction FPs). Concentrated real-bug cluster found: ab-731-d1 has 5 Qs referencing legacy GPT-4 $0.03/1K + GPT-3.5-turbo $0.002/1K token pricing -> bug #22 logged.~~ — commits 4b739d6 + c7f7bde (003b7095 session)
  • ✅ ~~P3.2 (this session): deprecated_check.py scanner — HIGH (hard-deprecated: AzureRm cmdlets, classic portal, CLI v1, old preview APIs) / MEDIUM (renamed: Azure AD, AAD, gsutil, aws-sdk v2, MIP/AIP scanner) / LOW (context-historic: bq, Azure classic). 0 HIGH hits portfolio-wide -- bank quality on hard-deprecated content is CLEAN. 225 MEDIUM mostly Azure AD terminology rename (70 Qs / 48 cert-files, spread across MS + vendor-integration certs). aws_sdk_v1_pkg regex calibrated mid-session after @aws-sdk/* v3-modular FP found (13 FPs eliminated). Bug #23 logged with Sush-decision proposal.~~ — commits 4b739d6 + c7f7bde (003b7095 session)
  • ✅ ~~#23 Entra rename SHIPPED (Path B selective MS-owned): 14 files / 13 Qs / 37 renames. 2-pass audit caught + reverted 15 of 28 initially-renamed Qs (12 using 'Azure AD' as intentional historical name in 'formerly/legacy' patterns; 3 using 'Microsoft Entra ID (Azure AD)' parenthetical alias).~~ — commit 42f4e37 (003b7095-step1)
  • ✅ ~~#21 Tier 1 + Tier 3 top-10 dead-link sweep SHIPPED (Path A'): 26 files / 441 learnLink edits. 8 manual overrides + 36 auto-chain to verified product hubs with HUB_ALLOWLIST + TOO_GENERIC guardrails. 3 Cisco residuals logged as bug #24.~~ — commit 50e4768 (003b7095-step2)
  • ✅ ~~#22 ab-731-d1 license-vs-credits rewrite SHIPPED (Path C, Voice Rule answer (c)): 5 Qs (d1-019, d1-020, d1-021, d1-022, d1-s06) pivoted from stale March-2023 GPT-4 token pricing to Build 2026 M365 Copilot architecture (per-user $30 + Copilot Credits $0.01/credit + Agent 365 governance). Rubber-duck caught 2 blocking issues (E7 out of d1-s06 numeric math; Credits never positioned as substitute for human-Copilot licence) + 8 non-blocking critiques, all adopted. Cert-identity error corrected (ab-731 is 'AI Transformation Leader' not 'Copilot Sales Specialist'). Zero new mashup HIGH; portfolio MEDIUM dropped 6781→6780. All 5 learnLinks validated 200.~~ — commits e873236 + cadb7b0 (b5cb9fa4 session)

Per-cert workflow (canonical, post-session-2)

For each cert in P2: 1. Run python portfolio_mashup_scan_v2.py (latest scanner) filtered to cert → fresh finding list 2. Triage each HIGH finding by opening the Q (use triage_q.py <cert> <qid> or batch_triage.py <cert> HIGH) 3. Classify: REAL_BUG vs FALSE_POSITIVE (per the rules below) 4. For REAL_BUG: fix via edit tool (sacred-field rewrite if needed) — OR if it's H4 enrichment-block contamination, apply B-Lite strip pattern (wipe whyWrong, delete hint + examTip; keep Q + correct + explanation) 5. Launch per-cert SME audit (research sub-agent) — read 5 random Qs + the rewritten Qs, validate against authoritative docs (use the learnLink field as source pointer). Capture any non-mashup bugs surfaced to discovered_bugs SQL table. 6. Apply SME-flagged HIGH+MEDIUM fixes 7. Build clean: python -m json.tool each modified file → npm run build → verify dist artifact has expected Q count (rubber-duck rule: build silently skips invalid JSON) 8. Commit with EXPLICIT paths (parallel-safe git rule, never git add .) 9. git pull --rebase && git push 10. SLA smoke: az-900 + checkout + practice page + per-cert questions.json + per-cert practice page 11. Update certs SQL: status='done', high_fixed=N, high_fp=M, notes=...

Triage rules (proven from sessions 1+2)

H1 (whyWrong has key for correct option): - Always real. If text reinforces correct → move to explanation. If text contradicts stem → mashup evidence. If meta-commentary → delete. - TF Qs with whyWrong[true] when correct=true: replace with whyWrong[false] explaining why False is wrong.

H2 (whyWrong text contains "is the correct/right"): - ~95% are false positives. Apply rule: "Does whyWrong claim THIS wrong option is correct (real mashup) OR does it correctly contrast the wrong option against the right answer (FP)?" - FP categories to record: contrast-to-right-answer, correct-in-other-context, correct-behavior-not-correct-option, partial-credit-not-complete

H4 (enrichment-block divergence, ratio ≤ 0.07): - ~100% real inside juniper, ~60% real cross-cert - Pattern: Q + correct + explanation on topic A; whyWrong + hint + examTip on completely different topic B - Fix: B-Lite strip (whyWrong: {}, delete hint + examTip). Keep Q + correct + explanation. Re-author later.

H3 MEDIUM (token overlap heuristic): - ~95% FP on scenario-based Qs where stem uses high-level vocab and correct option has technical specifics - Triage by semantic incoherence, not token overlap

H2 triage shortcut (calibrated 2026-06-03 sessions c0952d49 + 1ef77611 over 165 Qs / 18 certs): - 96.4% FP rate confirmed. Sush's authoring style produces "contrast-to-right-answer" whyWrong text that triggers H2 systematically. - FP signature: whyWrong[X] (where X is a WRONG option) contains phrases like "the right tool", "the correct approach", "the proper way", "is indeed the correct" — and the sentence is describing what the correct option does, by contrast. - REAL signature: whyWrong[X] claims the wrong option itself is correct, OR options/correct/explanation tell different stories (corrupt_options mashup — same class as juniper-jncis-ent-d2-008 + mb-800 d4-034/d4-035). - Fast path: when batch-triaging a cert with all-H2 hits, read the EXPLANATION first. If explanation supports the marked correct option and whyWrong sentences contrast WRONG-vs-CORRECT, it's an FP. If explanation supports a DIFFERENT option than marked correct, or describes a different topic than Q stem, it's a corrupt_options mashup — delete.

Corrupt-options mashup (real bug class at the H1/H2 layer — DELETE): - Signature: Q stem says topic A; options are about topic B (or generic admin pages); correct marker points at something that doesn't match either; explanation describes the right answer but it's NOT in the options - Examples: juniper-jncis-ent-d2-008 (STP Q with BGP options), mb800-d4-034 (customer return Q with admin-page options), mb800-d4-035 (vendor credit Q with inventory journal explanation) - Fix: DELETE the Q. Cannot be repaired without inventing content (which would need Sush's voice approval per Voice Rule). Under-representation > misleading content per Phase D rule.

Marker-swap mashup (real bug class — REPAIR by marker swap; calibrated 2026-06-03 session 1ef77611): - Signature: whyWrong[X] (where X NOT in correct) contains VERBATIM CAPS "This IS the correct" or "This IS the right" — author-artifact signal that the option marker was changed without updating whyWrong, leaving the now-wrong option's whyWrong text claiming it IS the right answer - Discriminator: capital "IS" near sentence start (case-sensitive). Lowercase "this is the correct" is the standard H2 contrast pattern (FP — 6 portfolio hits all confirmed FP) - Examples: pl300-d1-012, d1-013, d1-016, d1-021 (all repaired in 49f0cf6) - Often paired with off-topic enrichment — all 4 pl-300 marker-swap Qs ALSO had hint/examTip/realWorld copy-pasted from completely unrelated Qs (H4 didn't catch because vocabulary overlap kept ratio above 0.07 threshold) - Fix: REPAIR by swapping correct marker to the actual right answer (the right answer IS in the options — no content invented). Rewrite whyWrong for the now-wrong option. B-Lite strip off-topic hint/examTip/realWorld if present. NO Voice Rule blocker — author intent was already to mark the actual right answer; we're correcting a marker bookkeeping bug. - Detection: python marker_swap_check.py (portfolio-wide, ~5s). Also wired as HARD-FAIL pre-commit hook in inject_phase_b.py.


Discovery infrastructure — capture EVERYTHING

You will find bugs that don't fit the mashup class. Capture them, don't forget them.

Persistence model (cross-session)

The session SQLite is wiped on every new session. The persistent store is a git-tracked JSON file: - Location: C:\ssClawy\guided\files\discovered_bugs.json - Shape: array of bug objects matching the SQL schema below - Lifecycle: - Session start: read JSON → load into session SQLite discovered_bugs table via SQL INSERTs - Work in session: insert + update via SQL (SQLite is the fast queryable working copy) - Session end: export SQL table back to JSON → commit JSON with explicit path

Helper scripts

In C:\ssClawy\guided\scripts\bugs\ (create if missing): - load_bugs.py — reads discovered_bugs.json → emits SQL INSERT statements (run at session start) - save_bugs.py — reads SQLite discovered_bugs → writes JSON (run at session end)

If those scripts don't exist yet, create them — first session that needs cross-session capture builds the tooling.

Schema (session SQLite + JSON shape)

CREATE TABLE discovered_bugs (
  id INTEGER PRIMARY KEY AUTOINCREMENT,
  cert TEXT NOT NULL,
  qid TEXT,                        -- nullable for cert-level issues
  bug_class TEXT NOT NULL,         -- e.g. 'factual_error', 'deprecated_ref', 'typographic_artifact', 'corrupt_options', 'enrichment_stripped'
  severity TEXT,                   -- HIGH / MEDIUM / LOW
  source TEXT,                     -- 'h1_scanner', 'h2_scanner', 'h4_scanner', 'sme_audit', 'manual_triage', 'sush_flag'
  discovered_session TEXT,         -- session ID
  discovered_at TEXT,              -- ISO date
  description TEXT,                -- what's wrong
  proposed_fix TEXT,               -- how to fix
  status TEXT DEFAULT 'open',      -- open / in_progress / fixed / wontfix / duplicate
  fixed_session TEXT,
  fix_commit TEXT,
  fixed_at TEXT
);

When to insert

  • SME audit flags a factual error (vault d4-008 class) → insert immediately
  • You spot a typo / deprecated reference / pricing claim during triage → insert
  • You find a Q with no good fix in current scope → insert with status='open'
  • A bug class repeats 2+ times across certs → insert + propose building a scanner for it

When to query

  • Start of every session: read JSON + run SELECT * FROM discovered_bugs WHERE status='open' ORDER BY severity, discovered_at — surfaces the queue
  • After each cert triage: check if any open bugs relate to the cert you just worked on → bundle the fix

Reading the queue at session start

SELECT id, cert, qid, bug_class, severity, description
FROM discovered_bugs
WHERE status = 'open'
ORDER BY
  CASE severity WHEN 'HIGH' THEN 1 WHEN 'MEDIUM' THEN 2 WHEN 'LOW' THEN 3 ELSE 4 END,
  discovered_at;

Self-learning protocol — improve the system as you go

Rule: when you see a bug class 2+ times, build a scanner for it

The portfolio has 31,212 Qs. Manual inspection cannot find systemic bugs. Each scanner is a force multiplier.

When you discover a new bug class (e.g., SME catches "Q references deprecated az vm extension set syntax"): 1. Note the pattern + its detection signal (e.g., regex matching deprecated CLI syntax) 2. If you see it AGAIN in another cert during the same session OR a previous session: build a scanner 3. Scanner template: extend portfolio_mashup_scan_v2.py or create <bug_class>_check.py in C:\ssClawy\guided\ 4. Calibrate per Rule #6 (data-first sequence): probe 1000+ honest Qs to find natural threshold cliff before deciding HIGH/MEDIUM/LOW 5. Run portfolio-wide → log all hits to discovered_bugs 6. Triage + fix the real bugs

Rule: when heuristic over-fires, recalibrate

If you triage 20+ findings and ≥80% are FPs, the heuristic is too loose. Either: - Tighten the threshold (use the data — check ratio distribution of confirmed-real vs confirmed-FP cases) - Reclassify (HIGH → MEDIUM, MEDIUM → advisory) so triage focuses on high-precision signals first - Add a secondary filter that eliminates the FP class (e.g., "exclude H2 hits where wrong-option-whyWrong references correct-option text")

Rule: workflow friction → propose an update

If a step in the per-cert workflow takes >3× the expected time or causes a real mistake, propose a workflow update at end of session. Update THIS file before committing it. Future sessions inherit the improvement.

Rule: SME audit is your secondary scanner

Per kickoff doc, SME audit is mandatory after rewrites. It also catches all bug classes no heuristic detects (factual errors, deprecated refs, soft leaks, cross-doc inconsistencies). Run SME audit liberally during P2 triage — capture EVERYTHING it surfaces to discovered_bugs.


Tooling (in repo, ready to use)

  • C:\ssClawy\guided\portfolio_mashup_scan_v2.py — mashup scanner (H1/H2/H3/H4). P10 update (2026-06-03 session 7bb2af7a): H2 reclassified HIGH→MEDIUM based on 320-hit / 97-cert calibration (97.5% FP).
  • C:\ssClawy\guided\portfolio_mashup_scan.py — v1 deprecated, kept for reference
  • C:\ssClawy\guided\template_boilerplate_check.py — Phase B template-boilerplate scanner (B1-B12 phrases). Probe-only; supports --cert <slug>, --json <path>, --strict-fail for CI. B7-B12 added 2026-06-03 session c0952d49 for the az-104 secondary cluster. B12 uses qstem_only=True flag — only scans question field with end-anchored regex.
  • C:\ssClawy\guided\strip_template_boilerplate.py — Phase B template-boilerplate stripper. Pairs with the scanner. Modes: --dry-run, --apply, --cert <slug>, --sample <N>. Two code paths: strip_field() for B1-B11 (with trailing-period restoration) and strip_question_stem() for B12 (no period restoration; question stems already have their own terminator). Preserves JSON formatting matching b_lite_strip.py.
  • C:\ssClawy\guided\inject_phase_b.py — Phase B authoring tool. Pre-flight gates: schema check + banned-phrase scan + mashup heuristic (cc01005) + template-boilerplate heuristic B1-B12 (49a668d auto-extended by 80bc25c) + marker-swap heuristic (ea752df). --check-only for CI; --strict-mashup for H2 hard-fail (now opt-in maximum strictness since H2 is MEDIUM by default post-P10).
  • C:\ssClawy\guided\marker_swap_check.py — marker-swap scanner (whyWrong[X] for wrong option claims it IS correct). Case-sensitive verbatim CAPS "This IS the correct/right". 100% real-bug rate on calibration sample. Probe-only; --cert <slug>, --json <path>, --strict-fail for CI. Also imported by inject_phase_b.py as a HARD-FAIL gate.
  • C:\ssClawy\guided\learnlink_validator.py — async portfolio dead-link scanner (NEW 2026-06-03 session 7bb2af7a). Walks all learnLink fields (top-level + subQuestions), dedupes, HEAD-checks (retries GET on 403/405) with 30-concurrent + 30s-timeout + 2 retries. Classifies: ok / redirect_minor / redirect_major / dead / server_error / client_error / network_error / unknown_error. Args: --sample N / --cert <slug> / --out-json / --out-md. Exit code 1 if any dead links found (CI-friendly). Periodic scanner (NOT a pre-commit gate — link liveness drifts independently of authoring). Recommended monthly run.
  • C:\ssClawy\guided\pricing_check.py — portfolio pricing-claim scanner (NEW 2026-06-03 session 003b7095, P3). HIGH (dollar+per-unit context within ~45 chars: $30/user/month, $1.50/GB/month, $140 USD per month, $0.01 USD per hour) / MEDIUM (bare dollar amounts, often scenario fiction) / LOW (free_tier + savings_plan terms). Args: --cert <slug>, --severity HIGH|MEDIUM|LOW, --json <path>, --md <path>, --no-scenario (skip scenario field), --strict-fail. Calibrated 40% TP rate on 25-Q diverse triage sample (60% FPs are business-scenario fictional cost amounts). Periodic-only scanner; recommended quarterly. Baseline scan at files/pricing-scan-full.{json,md}.
  • C:\ssClawy\guided\deprecated_check.py — portfolio deprecated-pattern scanner (NEW 2026-06-03 session 003b7095, P3). HIGH (hard-deprecated: AzureRm cmdlets, classic portal URLs, Azure CLI v1, old preview APIs, AWS SDK v1 explicit refs) / MEDIUM (renamed: Azure AD → Microsoft Entra ID, AAD, gsutil, aws-sdk v2 package, MIP/AIP scanner) / LOW (bq CLI, Azure classic refs). Same args as pricing_check. 0 HIGH portfolio-wide (bank quality clean!), 225 MEDIUM mostly Azure AD rename cluster. aws_sdk_v1_pkg regex calibrated mid-session to exclude @aws-sdk/* v3 modular packages ((?<![@/\w-])aws-sdk(?![-/\w])). Recommended quarterly run paired with pricing_check.py + learnlink_validator.py.
  • C:\ssClawy\guided\leak_check.py — soft answer leak heuristic
  • C:\ssClawy\guided\test-guided-qa.cjs — full Playwright QA suite (MANDATORY before pushing PracticeQuiz changes; OPTIONAL for JSON-only changes)
  • C:\ssClawy\guided\test-banned-phrases.cjs — 17 specific template-engine artifact patterns
  • C:\ssClawy\guided\hugo-safe.ps1 (in aguidetocloud-revamp) — NOT relevant here (guided uses Astro)

Session-state tooling (reusable)

Session 9e5cc23f files: - files/triage_q.py <cert> <qid> — single-Q triage view - files/batch_triage.py <cert> HIGH|MEDIUM — compact batch view - files/all_h1.py — H1 portfolio dump - files/gsc-cert-top20.py — GSC traffic ranking - files/gsc-cert-top30.json — cached GSC top-30 cert ranking - files/load_certs.sql — re-load tracker table on next session (126 certs)

Session 89327d9c files: - files/h4_calibrate.py — H4 calibration probe (reuse for new scanners) - files/inspect_q.py — single-Q dump - files/b_lite_strip.py — B-Lite stripper (reusable for any cert with H4 contamination) - files/validate_json.py — JSON parse validator - files/scan-final/ — pre-fix v2 scan baseline - files/scan-post-fix/ — post-91fe2cb v2 scan - files/phase-d-scanner-findings.md — H4 discovery report - files/phase-d-session2-handoff.md — session 2 full handoff (read if you need deep detail)

Session 7bb2af7a files (P10/P11/P12): - files/p11_probe.py — calibration probe template (REUSE for any new heuristic): 5 metrics × 3 populations (REAL/H2FP/HONEST) cliff analysis. Demonstrated that no token-overlap metric discriminates whyWrong-vs-Qstem drift; pattern is reusable for future scanner candidates per Rule #6 data-first sequence. - files/p11_probe_data.json — raw drift-metric data (5 REAL + 50 H2FP + 200 HONEST samples). Future session: do NOT re-invent token-overlap drift detection — see this data for why. - files/prefix/ — pre-fix versions (via git cat-file blob) of pl-300-d1 + ab-731-d2 used to reconstruct known-real drift cases for calibration. - files/learnlink_probe.py — learnLink field-shape probe (Rule #6 pre-build probe for P12). - files/learnlink-results-full.json — full portfolio scan output (7,153 URLs). - files/learnlink-summary-full.md — human-readable findings for Sush strategic call.

Session 003b7095 files (P3 + bug 21 prioritization): - files/pricing_probe.py — Rule #6 data-first probe template for the pricing-claim scanner. Reusable as a starting point for ANY new heuristic candidate. - files/deprecated_probe.py — same probe template for deprecation patterns. - files/find_legacy_gpt_pricing.py — bug #22 reproduction script. Finds all Qs portfolio-wide with legacy GPT-4 / GPT-3.5-turbo per-1K token pricing patterns. Re-run if Sush picks a refresh path. - files/pricing_triage_sample.py — diverse-cert HIGH-hit sampler (picks 1 random Q per top cert) for scanner calibration triage. - files/learnlink_priority_rank.py — cross-ranks dead URLs vs GSC top-30 traffic for bug #21 decision support. Re-run after any GSC top-30 update to refresh tier rankings. - files/pricing-probe-results.json + files/pricing-scan-full.{json,md} — pricing scanner baselines (rerun monthly for drift). - files/deprecated-probe-results.json + files/deprecated-scan-full.{json,md} — deprecated scanner baselines. - files/legacy-gpt-pricing-hits.json — 5-Q stale GPT-4 pricing cluster (bug #22 evidence). - files/learnlink-priority-ranking.{json,md} — dead-link tier ranking (Tier 1 = GSC top-10 actionable / Tier 4 = zero-traffic deferrable). - files/apply_entra_rename.py — selective Azure AD -> Microsoft Entra rename engine (--dry-run/--apply/--cert). Includes longest-first ordering for compound names (Connect/Sync/B2C/B2B/PIM/etc.) - files/audit_rename_safety.py — first-pass safety detector for 'formerly Azure AD' / 'legacy' / 'older name' historical-use patterns. - files/audit_alias_parens.py — second-pass detector for 'Microsoft Entra ID (Azure AD)' parenthetical-alias patterns. - files/classify_azure_ad_v2.py — correct word-boundary classifier (replaces v1 which had 'azure ad' substring FP matching 'Azure Advisor' / 'Azure Application Gateway'; lesson archived in session journal). - files/revert_flagged_qs.py — field-level revert helper (uses git show HEAD:path to restore individual Qs to pre-rename state). - files/prep_tier1_worklist.py — builds Tier 1 + Tier 3 worklist from priority ranking. - files/replace_dead_links.py — async HEAD-verified dead-link replacer with HUB_ALLOWLIST + TOO_GENERIC guardrails + MANUAL_OVERRIDES table. Reusable for future link drift cycles. - files/inspect_replacements.py — manual review helper for replacement quality. - files/tier1-tier3-worklist.json — session input data. - files/tier1-tier3-replacements.json — full replacement log (44/47 resolved, 3 Cisco residuals). - files/update_bugs_post_ship.py — bug record updater for #21, #22 (enriched with Build 2026 context), #23, #24 NEW.


Stop criteria — when to ping Sush via ask_user

Per the kickoff: "Don't ping me unless C6 SME surfaces something needing my call."

Translate to: - SME audit surfaces a content scope question (e.g., "should this Q be rewritten or deleted entirely?") - Architectural choice with significant cost (e.g., a new scanner type would take >8h to build, or a fix pattern would touch >50 certs) - Practice exam SLA smoke fails post-deploy and revert isn't trivial - A bug class repeats so often it changes the program's scope (e.g., "factual errors are 10× more common than mashup bugs — should we pivot Phase D into Phase F factual-correctness?") - Per-cert effort blows up 3× the estimate - Sush-voice content question (any customer-facing rewrite needs his voice approval per Voice Rule)

Otherwise: full autonomy. Log everything to discovered_bugs + session journal + this doc.


Session end checklist

  1. Update certs SQL: mark each cert worked status='done' / 'high_done' / 'blocked'
  2. Update discovered_bugs SQL: any new bugs found go in open; bugs you fixed go to fixed with fix_commit
  3. Update THIS doc:
  4. Bump the "Last updated" line at top
  5. Move completed P-items from "What's open" into "What's shipped"
  6. Add any new P-items discovered
  7. Tighten the hours estimate if reality differed
  8. Update ~/.copilot/session-journal.md with per-session entry
  9. SLA smoke green (3 curls)
  10. Commit + push + git pull --rebase

Resume one-liner (give this to Sush for next session)

Hey Atlas — start the Phase D pricing sweep follow-on, top-GSC certs first.

(or the generic continue trigger if you want Atlas to pick: Hey Atlas — continue Phase D. Full autonomy. Stop-criteria only.)


  • Original kickoff: C:\ssClawy\guided\files\portfolio-mashup-cleanup-kickoff.md
  • Session 1 handoff (top-20 triage + 5 mashup fixes): ~/.copilot/session-state/9e5cc23f-37cd-4b5e-b796-29e715835fb4/files/phase-d-session1-handoff.md
  • Session 2 handoff (H4 scanner + juniper B-Lite): ~/.copilot/session-state/89327d9c-7309-438d-a8f6-3f257a278c82/files/phase-d-session2-handoff.md
  • Session 2 scanner findings (H4 calibration + cross-cert discovery): ~/.copilot/session-state/89327d9c-7309-438d-a8f6-3f257a278c82/files/phase-d-scanner-findings.md
  • Session 3 (7a664e44) template-boilerplate decision doc + scan: ~/.copilot/session-state/7a664e44-92eb-44f6-904e-89078e4d33ab/files/template-boilerplate-decision.md + template-boilerplate-scan.json
  • mb-800 Phase D deferred (12 sacred-needed Qs from before Phase D portfolio scope): C:\ssClawy\guided\files\mb-800-phase-d-followup.md