Files
veripath/docs/operations/task-reliability-ledger.md
T

77 lines
3.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Task Reliability Ledger
**Purpose:** Make the ≥100 error-free executions rule and Leonard scoring **measurable**.
Until a task type has documented runs, human approval remains required before any client-facing publish for that task type.
**Error definition (default):** Output that invents facts/numbers, mis-assigns Evidence Tier, includes solution language in Layer 1/2, omits required Snapshot Date / recency, violates template structure, or would be unsafe to show a client without correction.
---
## Publish-gate counters (automation eligibility)
| task_type | runs | errors | last_error_date | notes |
|-----------|------|--------|----------------|-------|
| gbp_snapshot_ingest_r1_r3 | 1 | 0 | — | First scored run 2026-08-03 DS4FLASH0731; 9/10 PASS; minor phone redaction + template shape |
| data_inventory_draft_l1a | 1 | 0 | — | Overcome Fitness public baseline written 2026-08-03 |
| threat_register_draft_l2 | 1 | 0 | — | Overcome Fitness; T-002 updated + T-011 added via R1/R2 |
| review_reply_draft | 0 | 0 | — | Not auto-publish |
| gbp_field_edit_draft | 0 | 0 | — | Not auto-publish |
| qa_answer_draft | 0 | 0 | — | Not auto-publish |
**Rule:** A task_type may bypass human approval **only after** ≥100 completed runs with **zero errors** for that exact task_type, documented here. Reliability does not transfer across task types.
---
## Leonard / agent scored runs (rubric 02 × 5 criteria)
| date | agent | client | task_type | score_total | zeros? | pass? | reviewer | notes |
|------|-------|--------|-----------|-------------|--------|-------|----------|-------|
| 2026-08-03 | DS4FLASH0731 (Leonard caveman) | Overcome Fitness | gbp_snapshot_ingest_r1_r3 | 9/10 | No | Yes | Director (Grok) | R1+R2 correct; R3 correctly not fired. Phone partially redacted (not in intake). Flat table vs inventory Notes format (template 1/2). Severity “Standard” on completeness vs register Major. |
**Pass:** no zeros and total ≥ 7/10.
**Rubric detail (this run)**
Evidence 2 · Tier 2 · Scope 2 · Recency 2 · Template 1
---
## Process log (first real client)
### 2026-08-03 — Overcome Fitness (first real non-Phoenix cold audit)
**What happened**
1. Human filled GBP observations (Maps) for Overcome Fitness.
2. Intake corrected live (booking button is on website, not GBP; services = strength training only; Updates/Posts present but empty; claimed status Unknown).
3. GBP snapshot form + intake template + Leonard execution card upgraded to **v1.1** so those fields are captured next time without chat back-and-forth.
4. Data Inventory GBP row written (Tier 2, Snapshot Date 2026-08-03).
5. Threat Register updated: T-002 evidence refreshed (R1); **T-011** added (R2 — unclear claimed + thin reviews).
6. Client README updated. Status remains **Draft Benchmark** (not Locked).
7. **First scored Leonard run:** DS4FLASH0731, inlined execution rules (no repo file access), caveman mode. **9/10 PASS.** Logged above.
**Agents involved**
- Director (Grok) led mapping, commits, scoring.
- Leonard = DS4FLASH0731 (first scored ingest).
**Artifacts**
- `docs/clients/overcome-fitness/data-inventory-v1.md`
- `docs/clients/overcome-fitness/threat-register-v1.md`
- `docs/clients/overcome-fitness/README.md`
- `tools/gbp-snapshot-form.html` (v1.1)
- `docs/templates/gbp-snapshot-intake.md` (v1.1)
- `docs/agents/leonard-gbp-execution-card.md` (v1.1)
**Next for this client**
- Human accept Draft Benchmark or request edits.
- Optional: re-run on RTX 3060 target model with same intake; compare scores.
- Optional: deeper social/website surfaces.
---
## How to log a run
1. After human review of an agent draft, add one row to the appropriate table.
2. If the draft required material correction for an error (see definition), increment `errors` and set `last_error_date`.
3. Do not count unscored exploratory chat as a run.
**Last updated:** 2026-08-03