72 lines
3.4 KiB
Markdown
72 lines
3.4 KiB
Markdown
# Task Reliability Ledger
|
||
|
||
**Purpose:** Make the ≥100 error-free executions rule and Leonard scoring **measurable**.
|
||
Until a task type has documented runs, human approval remains required before any client-facing publish for that task type.
|
||
|
||
**Error definition (default):** Output that invents facts/numbers, mis-assigns Evidence Tier, includes solution language in Layer 1/2, omits required Snapshot Date / recency, violates template structure, or would be unsafe to show a client without correction.
|
||
|
||
---
|
||
|
||
## Publish-gate counters (automation eligibility)
|
||
|
||
| task_type | runs | errors | last_error_date | notes |
|
||
|-----------|------|--------|----------------|-------|
|
||
| gbp_snapshot_ingest_r1_r3 | 0 | 0 | — | Leonard execution card only; form/template/card at v1.1 as of 2026-08-03 |
|
||
| data_inventory_draft_l1a | 1 | 0 | — | Overcome Fitness public baseline written 2026-08-03 |
|
||
| threat_register_draft_l2 | 1 | 0 | — | Overcome Fitness; T-002 updated + T-011 added via R1/R2 |
|
||
| review_reply_draft | 0 | 0 | — | Not auto-publish |
|
||
| gbp_field_edit_draft | 0 | 0 | — | Not auto-publish |
|
||
| qa_answer_draft | 0 | 0 | — | Not auto-publish |
|
||
|
||
**Rule:** A task_type may bypass human approval **only after** ≥100 completed runs with **zero errors** for that exact task_type, documented here. Reliability does not transfer across task types.
|
||
|
||
---
|
||
|
||
## Leonard / agent scored runs (rubric 0–2 × 5 criteria)
|
||
|
||
| date | agent | client | task_type | score_total | zeros? | pass? | reviewer | notes |
|
||
|------|-------|--------|-----------|-------------|--------|-------|----------|-------|
|
||
| — | — | — | — | — | — | — | — | No formal scored Leonard runs yet |
|
||
|
||
**Pass:** no zeros and total ≥ 7/10.
|
||
|
||
---
|
||
|
||
## Process log (first real client)
|
||
|
||
### 2026-08-03 — Overcome Fitness (first real non-Phoenix cold audit)
|
||
|
||
**What happened**
|
||
1. Human filled GBP observations (Maps) for Overcome Fitness.
|
||
2. Intake corrected live (booking button is on website, not GBP; services = strength training only; Updates/Posts present but empty; claimed status Unknown).
|
||
3. GBP snapshot form + intake template + Leonard execution card upgraded to **v1.1** so those fields are captured next time without chat back-and-forth.
|
||
4. Data Inventory GBP row written (Tier 2, Snapshot Date 2026-08-03).
|
||
5. Threat Register updated: T-002 evidence refreshed (R1); **T-011** added (R2 — unclear claimed + thin reviews).
|
||
6. Client README updated. Status remains **Draft Benchmark** (not Locked).
|
||
|
||
**Agents involved**
|
||
- Director (Grok) led mapping and commits following execution-card constraints.
|
||
- Leonard local model **not yet** run/scored on this intake.
|
||
|
||
**Artifacts**
|
||
- `docs/clients/overcome-fitness/data-inventory-v1.md`
|
||
- `docs/clients/overcome-fitness/threat-register-v1.md`
|
||
- `docs/clients/overcome-fitness/README.md`
|
||
- `tools/gbp-snapshot-form.html` (v1.1)
|
||
- `docs/templates/gbp-snapshot-intake.md` (v1.1)
|
||
- `docs/agents/leonard-gbp-execution-card.md` (v1.1)
|
||
|
||
**Next for this client**
|
||
- Human skim / accept Draft Benchmark or request edits.
|
||
- Optional: run same intake through Leonard (local model) and score vs rubric; log in scored-runs table.
|
||
|
||
---
|
||
|
||
## How to log a run
|
||
|
||
1. After human review of an agent draft, add one row to the appropriate table.
|
||
2. If the draft required material correction for an error (see definition), increment `errors` and set `last_error_date`.
|
||
3. Do not count unscored exploratory chat as a run.
|
||
|
||
**Last updated:** 2026-08-03
|