Files
veripath/docs/architecture/DOP-AI-Visibility-Unified-Framework-v1.md
T

328 lines
34 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# AI Visibility & Local Discovery — Unified Framework v1
## Convergence of the new Operating System draft into existing DOP governance
Status: Draft for review — resolves terminology collisions before this becomes canonical.
Supersedes: Nothing yet. This is the merge layer sitting between the new AI Visibility OS draft and the existing DOP repo (Agent Charter, Audit Playbook, Onboarding Manual, GLOSSARY, economic model).
Owner: Tony
Date: 2026-08-06
---
## 0. Why this document exists
Two frameworks now describe overlapping territory:
1. Existing DOP system — Layers 17 audit methodology (Layer 3 economics non-skippable), Tier 1 (owner-verified) / Tier 2 (public observation) evidence model, agent-detects/human-approves governance, filesystem-artifact-over-self-report verification standard.
2. New AI Visibility & Local Discovery OS — Stage 010 client lifecycle, Module AH workstreams, Verified/Indicative evidence classification, Get Found / Stay Found packaging.
Both are sound on their own terms. They were not built against each other, so they use similar-sounding words for different things. This document is the crosswalk — it decides which term wins, where, and why, so nothing downstream (reports, agent skills, client-facing language) inherits ambiguity.
---
## 1. Decision 1 — CORRECTED: there was no collision. Verified/Indicative already IS Tier 1/Tier 2.
Original assumption (wrong): I treated "Verified vs Indicative" and "Tier 1 vs Tier 2" as two different axes needing reconciliation, and proposed renaming the new OS's terms to "Observed/Inferred" to avoid a collision.
Ground truth, pulled directly from `audit-playbook-v1.md`:
> Evidence Model only — Every fact is Tier 1 (Verified) or Tier 2 (Indicative). No unmarked claims.
There is no second axis. Tier 1 is Verified. Tier 2 is Indicative. Same bucket, two labels used together, always. The playbook's Layer 1 rule reinforces this: "Public observation = Tier 2 until owner access or stronger verification upgrades it" — meaning Tier 1/2 already encodes *both* source authority and observation confidence in one scale, not two separate dimensions.
Resolution: Delete the "Observed/Inferred" rename — it solved a problem that doesn't exist. Module A's Verified/Indicative language (Section 11) is already correct as written and requires no change. Use "Tier 1 (Verified)" and "Tier 2 (Indicative)" together in all client-facing and internal documentation going forward, matching existing playbook style.
One real distinction worth preserving, not as a rename but as a note: the playbook's Tier 2 default rule — "Public observation = Tier 2 until owner access or stronger verification upgrades it" — is a *procedural* rule about how tiers get assigned, not a different taxonomy. Module A's Evidence Record schema (Section 20) should carry this default explicitly: any fact sourced from an AI engine output or public scrape starts at Tier 2 regardless of how confidently it was captured, and only moves to Tier 1 if the client confirms it directly.
Addendum — evidence authority nuance (added after Qwen review). The single Tier 1/2 scale holds, but different fact classes have different verifying authorities, and conflating them causes real confusion — "we verified the audit run happened" is not the same claim as "we verified the business fact inside it is true." Three classes:
1. Business-entity facts (name, phone, address, hours, services, categories, profile ownership). Default Tier 2 when observed publicly. Promotes to Tier 1 only via owner confirmation or owner-access verification.
2. Audit-run facts (prompt executed, engine captured, timestamp recorded, raw output stored, competitor appeared in captured output). Default Tier 1 when supported by stored filesystem/git/database artifacts — this verifies the audit *event* occurred, not that the business fact inside the output is true.
3. Public claims inside AI outputs (an engine stating the client's phone number, hours, or service scope). Default Tier 2 until owner confirmation or owner-access verification upgrades it.
Action: add this three-class breakdown to GLOSSARY under the Tier 1/2 entry, so "verified" is never ambiguous between "the run happened" and "the fact is true."
---
## 2. Decision 2 — CORRECTED: Layers are the whole engagement sequence, not audit depth inside Stage 2. Real gap found.
Original assumption (wrong): I guessed Layers 17 were a depth-scale nested inside Stage 2 (Discovery & Diagnosis).
RESOLVED — verified against `path-to-poc-sequencing.md`, pulled directly from main by Ty on 2026-08-06. The earlier "real gap found" claim was wrong. Layers 4 and 5 are fully specified in the actual sequence doc, each with entry criteria, exit criteria, required data, blocked-if-missing conditions, and a primary artifact — requirements-template.md (Layer 4) and workflows-template.md (Layer 5) both already exist in docs/templates/. This framework's earlier claim that these had "no equivalent" was built on an inability to reach the file, not on the file actually being empty. Confirmed sequence:
1. Account Intelligence — 1a Public baseline / 1b Owner-access. Artifact: data-inventory-v1.md
2. Threat Diagnosis. Artifact: threat-register-v1.md. Severity: Critical/Major/Minor
3. Economics & Pricing Validation — non-skippable, confirmed. Artifact: economics-v1.md / pricing-bands-v1.md
4. Response Requirements. Artifact: requirements-v1.md
5. Workflow Engineering. Artifact: workflows-v1.md
6. Agent Build & Test. Artifact: build-log-v1.md
7. POC Delivery & Retainer Conversation
Verified bridge map (no longer a working hypothesis):
| Client Stage | Layers (verified) | Notes |
|---|---|---|
| Stage 01 (Targeting & Lead Gen) | Pre-Layer 1 | ICP targeting, Free Check |
| Stage 2 (Discovery & Diagnosis) | Layer 1a + Layer 2 | Cold audit → Diagnostic Record |
| Stage 3 (Proposal & Packaging) | Layer 3 | Economics, non-skippable |
| Stage 4 (Onboarding & Access) | Layer 1b + Layer 4 | Owner-access verification *and* requirements — see structural correction below |
| Stage 5 (Remediation Sprint) | Layers 5 + 6 + 7 | Workflows, build/test, POC delivery |
| Stage 6 (Verification & Reset) | Re-run of Layers 12 tooling | Second diagnostic pass, not delivery |
| Stage 710 (Monitor → Offboard) | Post-PoC operations | Drift detection, stewardship |
Structural correction the sequence doc forces: Stage 4 cannot be Layer 4 alone. The sequence doc's own ordering — account intelligence → problems → requirements → workflows — plus the fact that owner-access verification (Layer 1b) has to happen before requirements can be finalized, means Stage 4 does double duty: owner-access work *and* requirements-gathering both land there. This also resolves where the Owner-Verified Correction Plan actually belongs (see Decision 7, corrected below) — it's Layer 1b content, not Layer 4.
Layer 7 still belongs at Stage 5, per the original catch — the sequence doc's own Layer 7 name ("POC Delivery") confirms this, it wasn't just a reasonable guess.
---
## 3. Decision 3 — Module A's Section 15 replaced with an enforceable Artifact A/B split
Non-Negotiable #5 states: *"No solution language in Layer 1 or 2. Diagnosis only. No 'fix this,' 'claim that,' or workflow suggestions."* Module A's original report structure (Section 15) bundled diagnostic content with solution content in one deliverable. Fix, converged on across two passes:
Artifact A — The Diagnostic Record (Layer 1 & 2 safe). Scope corrected after cross-check against existing repo templates — earlier drafts of this artifact duplicated data-inventory-v1.md (Layer 1) and threat-register-v1.md (Layer 2) instead of feeding them. Artifact A now holds only the measurement content with no existing home: Methodology & Limitations, the Core Visibility Index (70 pts quantitative, see below), Directional Context Indicators (Tier 2), Prompt-Level Evidence, and the per-engine sourcing table. Entity/signal state and evidence-ledger content are not restated here — they get transcribed into data-inventory-v1.md and threat-register-v1.md using those files' existing schemas as part of completing an audit run.
Artifact A is a source input to those two files, not a parallel record.
QA gate scope, refined after Qwen review — this was a real gap in the original rule. A blind banned-word scan across the whole document produces false positives: prompt text naturally contains "who should I call," raw engine output naturally contains "I recommend." The gate must apply only to authored diagnostic narrative, never to raw evidence payload:
| Field type | Gate applies? |
|---|---|
| Executive summary, observation narrative, competitor observation summaries | Yes — authored interpretation |
| Raw prompt text, raw engine output, source titles/URLs, platform field labels, quoted business names | No — observed evidence or system metadata |
Banned terms in the gated fields: should, recommend(s)/recommendation(s), fix/fixing, opportunity/opportunities, package(s), proposal/propose, solution(s), action plan, next step. Automated token matching is a first-pass control only — synonym evasion means human review stays required regardless.
Artifact B — The Strategy & Remediation Brief (Layer 3+). Scope corrected the same way. Earlier drafts duplicated requirements-v1.md (Layer 4) and workflows-v1.md (Layer 5) via a Prioritized Action List with Owner/Effort/Approval columns — that's workflow structure, already templated. Artifact B now holds only: the Opportunity Matrix (evidence × capacity × ticket size — the economic weighting layer with no existing home), Revenue-at-Risk scenarios, Package Recommendation, and Measurement Window. Actions themselves live in requirements-v1.md / workflows-v1.md; the Opportunity Matrix feeds into those by entry ID rather than restating them.
This enforces Non-Negotiable #5 by construction, not just by discipline — a scoped banned-word check on Artifact A's authored fields is trivial to automate as a QA gate, fitting the existing Gate pattern in Module A §19.
Severity vs. Priority — cross-reference, don't merge: the playbook's Threat Register uses Critical/Major/Minor (how bad is the problem — Layer 2, lives in Artifact A). Module A's Opportunity Prioritization Matrix uses Fix Now/Fix Soon/Monitor (how urgent is the fix — Layer 3+/5, lives in Artifact B). Keep both; note the distinction explicitly in GLOSSARY so a report never conflates a Critical *diagnosis* with a Fix Now *priority* as if they were the same field.
Scoring rubric issue — now resolved, not just flagged. Competitor Share Gap and Source/Citation Quality were qualitative bands dressed as part of a 0100 formula. Split into two separate things instead of forcing a single false-precision number:
- Core Visibility Index (quantitative, 70 pts, lives in Artifact A): Mention rate (35) + Mention strength/position (20) + Description accuracy (15). Everything here is directly countable from captured audit runs.
- Directional Context Indicators (Tier 2, reported separately, not folded into the 70): Source/citation quality, Competitor observation share, Engine variance. Client language: *"The Core Visibility Index is based on directly countable audit observations. The directional context indicators are qualitative and indicative, not a measured market share."* If a single number is still wanted for sales simplicity, call it the "Directional Visibility Index" and disclose in the methodology note that part of it is qualitative.
Action: Replace Section 15 in Module A's source doc with the Artifact A/B structure above, replace §10's scoring model with the Core Visibility Index / Directional Context Indicators split, and apply the field-scoped QA gate. Commit Artifact A's template under docs/agents/ (Layer 1/2 tooling, same category as the audit playbook) and Artifact B's template under docs/templates/remediation-brief-template.md — both now drafted in full, see files attached to this session.
---
## 4. Decision 4 — Module G inherits Leonard's actual failure record, not generic language
The gap: Module G ("Agent & Automation Layer") says "human-in-the-loop gates on all client-facing claims and live changes." True, but it doesn't cite *why* — it reads like a generic AI-safety boilerplate rather than a rule earned from real incidents.
Resolution: Module G should explicitly reference the standing protocol already established:
- Never trust agent self-reports of task completion — verify via git state or filesystem artifacts directly.
- This rule exists because of documented failure patterns: false self-completion reports, fabricated file content, fabricated SHA values (Leonard, prior sessions).
- The two-branch (develop/main) workflow was retired after a fabricated verification report; all commits now go directly to main with human verification substituting for the branch-gate.
Confirmed by the actual playbook — this isn't hypothetical, it's already implemented for one workflow. The GBP cold-audit pattern already enforces exactly this kind of constrained autonomy: *"Leonard maps fields and fires only R1R3"* via a fixed execution card, with an explicit ban on "free-form GBP reasoning." This is the working precedent for Module G's future agent skills (A1A7 in Module A) — each skill spec should follow the same pattern: a fixed execution card / trigger-rule set, not open-ended judgment, especially before a skill has a scored track record in the task reliability ledger.
Action: Rewrite Module G's human-gate language to cite this precedent instead of restating a generic principle. When Module A's Skills A2 (Audit Runner) and A3 (Mention Extractor) are actually built, they should follow the GBP execution-card pattern — fixed trigger rules, not free-form reasoning — until each has scored runs in the task reliability ledger.
---
## 5. Decision 5 — Section 4 (Data, Tools & Infrastructure) reflects the storage decision already made
The gap: The new OS's Section 4 lists data objects generically (client master record, prompt libraries, audit runs, etc.) as if the storage architecture were still undecided.
Resolution: It's decided. Replace the generic list with the actual hybrid architecture:
- Canonical business records — flat JSON in git (free version history, audit trail)
- Runtime check state — Postgres or SQLite (client slug, surface, field, canonical value, observed value, state, timestamps, raw response blob)
- Schema primitives — fact/surface separation with adapter-declared field mappings, a prohibition predicate type for ghost-listing detection (match = failure), type tags on every fact driving normalization at comparison time
Action: Replace Section 4's data-object list with this architecture description. Storage mapping now confirmed per-object rather than left as a general split:
| Data object | Primary storage | Secondary/export | Reason |
|---|---|---|---|
| Client master record | Git JSON | Postgres/SQLite cache | Canonical, versioned, auditable |
| Prompt library + version manifest | Git JSON | DB reference | Immutable audit trail |
| Competitor set | Git JSON | DB reference | Versioned scope |
| Audit run manifest | Git JSON | DB runtime record | Immutable metadata |
| Schema version | Git JSON / repo files | DB reference | Canonical versioned artifact |
| Task reliability ledger | Git JSON or DB (volume-dependent) | — | Needed before skill promotion (Decision 8) |
| Engine output record, prompt execution log | Postgres/SQLite | Audit run export | High-volume runtime data |
| Mention extraction, competitor mention, source citation records | Postgres/SQLite | Artifact A evidence ledger | Derived runtime data |
| Evidence record | Postgres/SQLite | Approved Artifact A export | Runtime ledger becomes approved artifact |
| Opportunity record | Postgres/SQLite | Approved Artifact B export | Prescriptive work belongs in Layer 3+ |
| Citation status ledger | Postgres/SQLite | Sprint report export | Runtime check state |
| GBP change log, drift detection event | Postgres/SQLite | Monthly stewardship export | Runtime change tracking |
---
## 6. Decision 6 — Packaging language checked against Revenue-at-Risk conservatism
Get Found / Stay Found / Custom packaging (Stage 3) is new relative to what's in the existing repo — first appearance. No conflict, but it should be reviewed once against the existing Revenue-at-Risk conservative-language rules before it's client-facing, specifically: package recommendation copy shouldn't imply guaranteed lift ("Get Found gets you found" reads differently than "Get Found addresses verified entity defects").
Action: Low priority — flag for a pass once packaging copy is actually drafted, not before.
---
## 7. Decision 7 — Merging Ty's business-readiness gap analysis (a second, parallel track)
Ty produced docs/DOP-BusinessReview_ty_grok_Aug062026.md — a business-readiness gap analysis, not a documentation-architecture reconciliation. It asks a different question than everything above: not "do our terms collide," but "what unmade decisions would surprise a client." Both tracks are needed; they hadn't been merged. This section merges them.
### 7.1 What Ty's document adds that this framework was missing
A. Per-engine sourcing breakdown — corrects Module A §8 (Engine Selection). Module A currently treats ChatGPT, Gemini, Perplexity, Claude, and Grok as symmetrical — same prompt set, same method, across all five. Ty's breakdown says they aren't:
Status: Tier 2 (Indicative)
Source: Ty business-readiness review, 2026-08-06
Validation required: direct audit observation or authoritative documentation before treating as operating fact.
| Engine | Primary local-answer source, working model | Highest-leverage signal, working model | Evidence tier |
|---|---|---|---|
| Gemini | Google's local graph | GBP, Maps/Places, reviews, categories, hours | Tier 2 |
| ChatGPT | Bing web index + partner place data + open web | Foursquare places data, Bing-visible consistency, business website, directory NAP | Tier 2 |
| Claude | Tools/APIs + web, often Google Places when looking up locals | Google Places/Maps-aligned data, clear website, consistent public facts | Tier 2 |
Action: Add this table, with its tier tag and source line intact, to Module A §8.3 (Engine Metadata Required) as context for why the *same* prompt run against different engines should not be scored as if testing the same underlying signal. This also means Skill A3 (Mention Extractor) needs engine-aware source attribution, not just a flat "which sources were cited" field. Resolved: the earlier flag that this table was an unmarked claim is now closed — it carries an explicit Tier 2 tag, source, and date, matching Non-Negotiable #1.
B. Two-phase owner pacing — CORRECTED mapping. Ty's Phase A (Correction Plan, heavy owner attention, done once) / Phase B (ongoing change management, light touch, owner notifies once) gives concrete shape to something the Path to PoC sequence names but doesn't fully template. Original mapping guessed Layer 4; verified mapping is different:
| Ty's phase | Maps to (verified) |
|---|---|
| Phase A — Correction Plan | Layer 1b (Owner-access verification) + Stage 4 — fills a real documentation gap in Layer 1b's exit criteria, not a Layer 4 gap (Layer 4 already has `requirements-v1.md`) |
| Phase B — Ongoing change management | Stage 7 (Ongoing Monitoring) content, reframed as *stewardship*, not just drift-watching |
This is a genuinely useful fill, just for a different layer than first proposed: Layer 1b's exit criteria imply canonical NAP, access matrix, no-change zones, and approval boundaries, but the sequence doc doesn't template that content. Ty's Correction Plan is a concrete candidate for that Layer 1b artifact.
C. "Human-heavy, automation-light V1" — a corrective to drift already visible in this framework. Ty's document names this explicitly as a scaling constraint. It reinforces, at the whole-service level, something Decision 4 only applied narrowly (the GBP execution-card pattern, no free-form agent reasoning). Module A's Skill A1A7 specs and the "Level 2 Agent-Assisted" language in the original OS draft imply more automation than a human-heavy V1 actually supports. Resolved — Module G replacement text:
## Module G — Agent & Automation Layer
### V1 posture
V1 is human-heavy and automation-light. Agent skills are tools underneath
human delivery, not autonomous operators. Skills A1A7 are Level 2 /
aspirational unless explicitly promoted through a dated decision and
supporting reliability evidence (see Decision 8 promotion gate).
### Standing verification controls
Never trust agent self-reports of task completion. Completion must be
verified through git state, filesystem artifact, stored raw output,
database row, hash/checksum where applicable, or human inspection of
the actual changed surface. Earned from documented failure patterns:
false self-completion reports, fabricated file content, fabricated
SHA values, verification claims unsupported by repository state.
### Execution-card pattern
Agent skills begin as constrained execution cards, not open-ended
reasoning tasks. Each card defines: trigger conditions, allowed inputs,
forbidden actions, deterministic output format, human approval gate,
evidence required to prove completion, rollback/escalation path.
Working precedent: the GBP cold-audit pattern — Leonard maps fields
and fires only fixed rules R1R3, no free-form GBP reasoning.
### Promotion rule
A skill receives expanded autonomy only after scored, successful runs
in the task reliability ledger (see Decision 8 thresholds). Until then:
no live client changes without human approval, no client-facing claims
without human review, no self-reported completion accepted as evidence.
D. Restraint as a design principle — genuinely new, not present anywhere else in either track. "Constant change can weaken confidence... leave alone what should stay stable" has no equivalent in Module A or the Path to PoC docs. Worth carrying into GLOSSARY or the eventual Phase B operating rules as an explicit principle: freshness is not the same as activity, and over-editing has a cost.
### 7.2 Exposure craft vs. content marketing boundary — resolved
The README lists V1 Scope Exclusions explicitly: Social media production, Paid advertising, Website redesign, Branding projects, Content marketing, Full SEO campaigns, CRM/email marketing. Ty's "exposure craft" sits close to that line. Boundary, now explicit:
In scope: platform category selection, attributes, service listing fields, structured short descriptions constrained by platform fields, Q&A seeded from owner-verified facts, schema properties derived from owner-verified data, canonical NAP enforcement, listing-field language matching customer search language.
Out of scope: open-ended copywriting, blog content, content marketing strategy, brand narrative development, social media production, paid advertising copy, website redesign copy.
Action: write this boundary into whatever doc formalizes Phase A / the Owner-Verified Correction Plan, so it isn't left to case-by-case judgment.
### 7.3 Consolidated open-decision list (merging both tracks)
Architecture / documentation track — resolved this pass:
- Confirm the Layer/Stage bridge map against the real `path-to-poc-sequencing.md` — done, verified, no longer a working hypothesis
- Split Module A §15 into Artifact A/B, commit both templates — done, both templates trimmed to remove duplication with existing repo templates and drafted in full
- Fix Module A §10's qualitative scoring bands — done, Core Visibility Index (70 pts quantitative) / Directional Context Indicators (Tier 2, separate) split
- Update GLOSSARY with Severity vs. Priority cross-reference — done, drafted in full at docs/GLOSSARY-additions.md
Business-readiness track (Ty's document) — still open:
- Website monitoring: exact always/never-checked list; remediation ownership; evidence standard
- V1 platform scope: which platforms are actually in v1; Foursquare/Bing Places depth
- Phase B intake channel, notice completeness requirements, turnaround SLAs, per-item-approval boundary
- AI sampling method, cadence, and owner-facing language — this is functionally a lightweight version of Module A's Free Check; worth explicitly naming as the same mechanism rather than a separate one
- Monthly report format when there's little to fix — "stewardship summary," not an empty repair list
- Retainer vs. project-only client criteria
Resolved this pass (was open, now closed):
- Tag the per-engine sourcing table with a source/date and Tier 2 marking — done
- Confirm Correction Plan's intended layer — done, re-homed to Layer 1b (owner-access verification), not Layer 4 — Layer 4's real deliverable is requirements-v1.md
- Add the exposure-craft / content-marketing boundary line — done, §7.2 above
- Add the Skills A1A7 = Level 2/aspirational note to Module G — done, full replacement text drafted above and in GLOSSARY
- Reconcile Artifact A/B against existing `data-inventory-v1.md` / `threat-register-v1.md` / `requirements-v1.md` / `workflows-v1.md` — done, both artifacts trimmed to non-duplicated content with explicit feed-forward sections
Still open from Decision 8:
- Observability variance/fallback policy for Module A — no owner, no target date yet
- Add the client-facing-surface restriction (Artifact A stays internal) to Module F and the onboarding manual as a non-negotiable line
Action: This section (7.3) is now the single master open-decision list. Future work should check items off here rather than re-deriving them from either source document separately.
---
## 8. Decision 8 — Internal rigor vs. delivery velocity (response to multi-source adversarial critique)
Three independent reviews converged on the same core tension without coordinating: process weight as a scaling and commercial risk. That level of convergence is itself the signal to resolve, not just note.
### 8.1 The critique, calibrated
Where it lands: the process-complexity-vs-velocity tension is real. A separate, genuinely new point that hadn't surfaced anywhere else: observability fragility — Module A's audit engine assumes reliable, repeatable capture across five consumer AI products indefinitely, with no fallback for vendor UI changes, geo-personalization drift, anti-bot friction, or rate-limiting on automated-looking usage. That's an operational risk, not a governance one, and governance language doesn't fix it.
Where it overstates the case, on reflection:
- It treated the business as a thin, specialized "AI Visibility" bet. The actual thesis (README) is broader — Silent Customer Loss prevention across Customer Path Integrity, AI Visibility, and Local Competitive Awareness. The AI-answer channel is one surface, not the whole bet.
- It read Artifact A's banned-word discipline as something clients would find evasive. Artifact A was never meant to reach a client raw — it's internal evidentiary backing behind a synthesized plain-language report. This needs to be made explicit and non-negotiable in the docs so it can't drift into a client-facing artifact by accident.
- It critiqued SMB unit economics against a volume-scale future that hasn't been chosen. The ICP matrix and "proof first, retainer after demonstrated value" motion already exist to prevent indiscriminate SMB volume-chasing at this stage. The critique is a legitimate warning for a later decision, not a diagnosis of the current one.
- The conservative evidence-language positioning was framed as accidental friction. It's a deliberate, already-stated bet (evidence discipline as differentiator vs. overpromising agencies) — untested, but not an oversight.
### 8.2 Resolution
- Internal apparatus stays mandatory at current PoC scale, no exceptions. Artifact A's banned-word gate, Tier model, execution cards, and the reliability ledger all remain required while the business is at 2 clients and proving the model. This is not the place to cut corners — it's also the phase with the least revenue pressure to justify cutting them.
- Client-facing surface is explicitly restricted and this boundary is now non-negotiable, not just a stated intent.
Clients see: the Correction Plan (Phase A), the prioritized action list, and stewardship reporting (Phase B). They never see Artifact A, Tier tags, execution-card mechanics, or Layer/Stage internals. Action: add a line to Module F (Client Delivery System) and to the onboarding manual stating this boundary explicitly, so it's enforced by documentation, not by individual judgment call.
- A dated, measurable promotion gate — made concrete, not left as "a decision is required": the playbook already has the machinery for this in its Training/Evaluation Rubric (02 per criterion, pass ≥7/10, **Benchmark quality = 910 with zero scope violations**). Use it directly as the velocity re-optimization trigger:
| Condition | Threshold |
|---|---|
| Scored runs logged in the task reliability ledger for a given skill | ≥5 runs |
| Runs scoring at Benchmark quality (910, zero scope violations) | ≥80% of those runs |
| Consecutive weeks with zero fabrication or self-report incidents for that skill | ≥4 weeks |
| Client count before *any* skill is eligible for promotion review | ≥3 (i.e., past pure PoC, not before) |
Only once a specific skill clears all four does it become eligible to move from execution-card-constrained to higher autonomy — decided per-skill, not as a blanket velocity switch for the whole system. This keeps the promotion mechanism itself evidence-gated, consistent with everything else in the repo, rather than a vague future intention.
- Observability fragility is logged as its own open operational item, separate from the complexity question: Module A needs an explicit variance and fallback policy — how much engine-to-engine variance is disclosed to clients vs. absorbed into the score, what happens when a vendor changes UI or rate-limits automated-looking sessions, and a defined minimum viable capture method if API access isn't available for a given engine. This has no owner yet and no target date. Action: add to §7.3's consolidated list as its own line, not bundled into the velocity gate above — a different kind of risk, different kind of fix.
Net effect: the critique doesn't change what gets built next, but it changes what "success" means before scaling — success now explicitly includes hitting the promotion thresholds above, not just delivering PoC results. That's a real answer to "when do we stop being this heavy," not just an acknowledgment that the question exists.
---
## 9. What's now resolved vs. still open
Resolved against the actual repo — audit-playbook-v1.md, README, path-to-poc-sequencing.md, and the four existing client/template files, all pulled from main (2026-08-06):
- No evidence-model collision existed. Verified = Tier 1, Indicative = Tier 2, one vocabulary, now with a three-class authority breakdown (business-entity facts / audit-run facts / public claims inside AI outputs).
- Layers 17 are the full, verified Path to PoC engagement sequence. The Stage↔Layer bridge map is confirmed, not a working hypothesis — including the Stage 4 double-duty correction (Layer 1b + Layer 4 both land there) that the sequence doc's own ordering forces.
- Layer 7/POC Delivery confirmed at Stage 5, not Stage 6 — the sequence doc's own naming confirms this was correct, not just a reasonable guess.
- Artifact A and Artifact B are trimmed to hold only content with no existing repo home, after discovering they originally duplicated data-inventory-v1.md, threat-register-v1.md, requirements-v1.md, and workflows-v1.md. Both now carry explicit Feed-Forward sections naming exact file paths, verified against live client directories (`docs/clients/overcome-fitness/`, `docs/clients/phoenix-salons/`).
- Module A's Section 15 replaced by the trimmed Artifact A/B split, satisfying Non-Negotiable #5 by construction, with the banned-word gate scoped to authored fields only (raw evidence exempt — avoids false positives on prompt text and engine output).
- The Core Visibility Index (70 pts quantitative) / Directional Context Indicators (Tier 2, separate) split resolves the earlier false-precision scoring issue.
- Module G's human-gate rationale ties to Leonard's documented failure patterns and the GBP execution-card precedent; full replacement text drafted.
- Data architecture and storage mapping (Decision 5) confirmed per-object.
- Ty's business-readiness gap analysis merged (Decision 7); per-engine sourcing table now carries an explicit Tier 2 tag and source date.
- The Owner-Verified Correction Plan is re-homed to Layer 1b, not Layer 4 — Layer 4's actual deliverable is requirements-v1.md.
- The multi-source complexity-vs-velocity critique resolved as Decision 8, including a measurable promotion gate against the existing reliability-ledger rubric.
Still open:
- Six business-readiness decisions from Ty's track (§7.3): website monitoring always/never list, V1 platform scope, Phase B intake/SLA rules, AI sampling cadence/language, monthly stewardship report format, retainer vs. project-only criteria.
- Observability fragility (Decision 8) — variance/fallback policy for Module A, no owner or target date yet.
- Requirements/workflows have templates but no live client instances yet — Artifact B's feed-forward targets (`requirements-v1.md`, `workflows-v1.md`) will need to be created per-client when Layer 4/5 work actually starts.
---
## 10. Immediate next step
The bridge-map verification that was blocking this document is done. Remaining work no longer has a single blocking dependency:
1. Commit the four files (this framework doc, Artifact A, Artifact B, GLOSSARY-additions) to main, direct commit per standing policy, no branch.
2. Work through §7.3's six remaining business-readiness decisions — none depend on further repo archaeology, they're just decisions to make.
3. Assign an owner and target date to the observability variance/fallback policy — still the only item in this document with no owner at all.
4. When Layer 4/5 work starts for a live client, create that client's requirements-v1.md and workflows-v1.md instances from the existing templates, and confirm Artifact B's Opportunity Matrix correctly feeds entries into them by ID.