Files
veripath/docs/agents/audit-playbook-v1.md
T

178 lines
7.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Audit Playbook v1.1
**Purpose:** Give any agent (especially Leonard) a fixed, Charter-compliant method for performing Layer 1 (Data Inventory) and Layer 2 (Threat Register) work.
**Source:** Extracted from the Overcome Fitness benchmark process (Director + Leonard iterations).
**Governing documents:**
- `docs/agents/agent-charter-v1.md`
- `docs/architecture/path-to-poc-sequencing.md`
- Templates in `docs/templates/`
**GBP-specific tooling:**
- Intake template: `docs/templates/gbp-snapshot-intake.md`
- Leonard execution card (minimal reasoning): `docs/agents/leonard-gbp-execution-card.md`
- HTML form (fill → copy/paste to Leonard): `tools/gbp-snapshot-form.html`
**Reliability / scores:** Record scored runs and publish-gate counters in `docs/operations/task-reliability-ledger.md`.
---
## Non-Negotiables
1. **Evidence Model only** — Every fact is Tier 1 (Verified) or Tier 2 (Indicative). No unmarked claims.
2. **Agents draft. Humans approve.** Nothing is locked without human sign-off.
3. **No invented data, numbers, or threats.**
4. **Order is fixed:** Data Inventory (Layer 1a, then 1b when available) before Threat Register (Layer 2). Never reverse.
5. **No solution language in Layer 1 or 2.** Diagnosis only. No "fix this," "claim that," or workflow suggestions.
6. **Stale data must be flagged.** If a snapshot is older than ~3060 days or its date is unknown, mark it and treat findings as provisional.
7. **Prefer the more constrained, evidence-bound position** when uncertain.
---
## Layer 1a — Public Data Inventory Rules (cold audit)
**Goal:** Record what is publicly observable, where it lives, and verification status. Nothing else.
### Required sources (minimum)
- Google Business Profile / GBP (or clear statement that it is missing/unclear)
- Website
- Primary social channel(s) visible without owner login
- Booking / scheduling surface as visible publicly (or explicit "not visible")
- Reviews / reputation surfaces if observable
### Required fields per source
- Access Status (public / pending owner / denied)
- Data Range Available (use date ranges, not vague "12 months")
- Snapshot Date
- Verification Tier (almost always Tier 2 for pure public observation)
- Notes / Gaps
- Snapshot Location path
### Rules
- Public observation = **Tier 2** until owner access or stronger verification upgrades it.
- Do not claim "owner access" unless confirmed in the current process.
- Missing sources are documented as gaps, not ignored.
- Technical findings (404s, schema, NAP conflicts, etc.) require a dated snapshot. Old findings must be marked for re-verification.
- **GBP without owner access:** use human snapshot via `tools/gbp-snapshot-form.html` (or markdown intake). Leonard maps via the execution card only — no free-form GBP reasoning.
### Exit criteria (human must confirm) — Layer 1a
- [ ] All accessible **public** sources recorded
- [ ] Every source has Tier + Snapshot Date
- [ ] Gaps explicitly listed
- [ ] No analysis or threat language present
- [ ] Human sign-off obtained
- [ ] Status set to Locked (1a) or Draft Benchmark with blockers listed
Cold audits may stop here. Do not claim "Layer 1 complete (engagement-ready)" without 1b or explicit acceptance of Tier 2 ceiling.
---
## Layer 1b — Owner-access Inventory (engaged)
**Goal:** Upgrade inventory with owner-gated sources; enable Tier 1 where justified.
### Additional rules
- Owner GBP / social / booking access documented per source
- History gaps remain gaps (no invented continuity)
- Exit only when human confirms owner-gated fields and tier upgrades
---
## Layer 2 — Threat Register Rules
**Goal:** Identify discrete, evidence-backed failures in customer acquisition or retention. Diagnosis only.
### Required fields per threat
- Threat ID (T-001, T-002…)
- Title (short, factual)
- Signal Source (must map to an inventory row)
- Evidence Snippet (specific, observable fact)
- Data Recency (must match or reference inventory snapshot dates)
- Estimated Severity (Critical / Major / Minor)
- Verification Tier
- Notes / Context (optional, still factual)
- Human Decision (Pending until review)
### Rules
- One primary fact per threat. Do not bundle multiple unrelated gaps under one ID.
- Evidence Snippet must be concrete (what was observed), not a general claim.
- Severity is preliminary and Tier-linked. Do not escalate severity beyond what the evidence supports.
- If data is stale, say so in Data Recency or Notes.
- No recommendations, fixes, or next actions.
- **GBP-related threats:** only create/update via execution-card trigger rules R1R3 when ingesting a GBP snapshot.
### Competitor evidence rules (mandatory)
- Competitor comparisons are **secondary context only**, never the primary claim of a threat.
- Any fact about a competitor's business (hours, rating, booking path, review count, etc.) is **Tier 2 by default**. Tier 1 is not available without verification from a source we control (almost never).
- Every competitor claim must include: **Snapshot Date** and **source URL or observation method**.
- A wrong claim about a third party is a **higher reputational risk** than a wrong claim about the client. Prefer omit over weak comparison.
- Client-facing comparative briefs must not treat competitor Tier 2 claims as proven superiority; use relative, provisional language only.
### Severity guidance (use conservatively)
- **Critical:** Directly blocks or severely impairs local discovery / conversion and evidence is clear.
- **Major:** Material friction or visibility gap with solid observation.
- **Minor:** Real but lower immediate impact, or evidence is thin.
When unsure, choose the lower severity.
### Exit criteria (human must confirm)
- [ ] Every threat traces to the inventory
- [ ] Every threat has Evidence Snippet + Data Recency + Tier
- [ ] No solution language
- [ ] No invented threats
- [ ] Competitor facts (if any) follow competitor evidence rules
- [ ] Human has recorded decisions (Verified / Downgraded / Rejected)
---
## Operating Pattern for Agents
1. Read the Charter and this Playbook first.
2. Work only inside the current layer.
3. Produce drafts that match the templates.
4. List every change made when revising an existing document.
5. Surface blockers and stale data explicitly.
6. Stop at the human review gate. Do not self-lock.
### GBP cold-audit pattern
1. Human fills `tools/gbp-snapshot-form.html` (or markdown intake).
2. Human pastes output to Leonard with: follow `docs/agents/leonard-gbp-execution-card.md` exactly.
3. Leonard maps fields and fires only R1R3.
4. Human reviews; Director may score against the rubric below and **records the score in the task reliability ledger**.
---
## Training / Evaluation Rubric (simple)
Score a Leonard (or any agent) run 02 on each:
| Criterion | 0 | 1 | 2 |
|-----------|---|---|---|
| Evidence grounding | Invented or vague | Mostly grounded | Fully concrete snippets |
| Tier discipline | Missing or wrong | Present but soft | Correct and consistent |
| Scope control | Drifted into solutions/requirements | Minor drift | Strict Layer 1/2 only |
| Recency honesty | Ignored stale data | Partially flagged | Explicitly flagged |
| Template fidelity | Structure broken | Mostly followed | Clean match |
**Pass threshold:** No zeros, and total ≥ 7/10.
**Benchmark quality:** Total 910 with zero scope violations.
Record each scored run in `docs/operations/task-reliability-ledger.md`.
---
## Current Benchmark Reference
- Client: Overcome Fitness
- Files: `docs/clients/overcome-fitness/data-inventory-v1.md` and `threat-register-v1.md`
- Status: Draft (Benchmark) — not yet Locked; GBP completeness still open; treat as **Layer 1a path**, not engagement-complete Layer 1
Future Leonard runs should be compared against this package and scored with the rubric above.
---
**Version:** v1.1
**Status:** Active for training
**Next upgrade trigger:** After first Locked client Layer 1a+2 package exists and at least one Leonard scored run is recorded in the ledger.