3.8 KiB
Task Reliability Ledger
Purpose: Make the ≥100 error-free executions rule and Leonard scoring measurable.
Until a task type has documented runs, human approval remains required before any client-facing publish for that task type.
Error definition (default): Output that invents facts/numbers, mis-assigns Evidence Tier, includes solution language in Layer 1/2, omits required Snapshot Date / recency, violates template structure, or would be unsafe to show a client without correction.
Publish-gate counters (automation eligibility)
| task_type | runs | errors | last_error_date | notes |
|---|---|---|---|---|
| gbp_snapshot_ingest_r1_r3 | 1 | 0 | — | First scored run 2026-08-03 DS4FLASH0731; 9/10 PASS; minor phone redaction + template shape |
| data_inventory_draft_l1a | 1 | 0 | — | Overcome Fitness public baseline written 2026-08-03 |
| threat_register_draft_l2 | 1 | 0 | — | Overcome Fitness; T-002 updated + T-011 added via R1/R2 |
| review_reply_draft | 0 | 0 | — | Not auto-publish |
| gbp_field_edit_draft | 0 | 0 | — | Not auto-publish |
| qa_answer_draft | 0 | 0 | — | Not auto-publish |
Rule: A task_type may bypass human approval only after ≥100 completed runs with zero errors for that exact task_type, documented here. Reliability does not transfer across task types.
Leonard / agent scored runs (rubric 0–2 × 5 criteria)
| date | agent | client | task_type | score_total | zeros? | pass? | reviewer | notes |
|---|---|---|---|---|---|---|---|---|
| 2026-08-03 | DS4FLASH0731 (Leonard caveman) | Overcome Fitness | gbp_snapshot_ingest_r1_r3 | 9/10 | No | Yes | Director (Grok) | R1+R2 correct; R3 correctly not fired. Phone partially redacted (not in intake). Flat table vs inventory Notes format (template 1/2). Severity “Standard” on completeness vs register Major. |
Pass: no zeros and total ≥ 7/10.
Rubric detail (this run)
Evidence 2 · Tier 2 · Scope 2 · Recency 2 · Template 1
Process log (first real client)
2026-08-03 — Overcome Fitness (first real non-Phoenix cold audit)
What happened
- Human filled GBP observations (Maps) for Overcome Fitness.
- Intake corrected live (booking button is on website, not GBP; services = strength training only; Updates/Posts present but empty; claimed status Unknown).
- GBP snapshot form + intake template + Leonard execution card upgraded to v1.1 so those fields are captured next time without chat back-and-forth.
- Data Inventory GBP row written (Tier 2, Snapshot Date 2026-08-03).
- Threat Register updated: T-002 evidence refreshed (R1); T-011 added (R2 — unclear claimed + thin reviews).
- Client README updated. Status remains Draft Benchmark (not Locked).
- First scored Leonard run: DS4FLASH0731, inlined execution rules (no repo file access), caveman mode. 9/10 PASS. Logged above.
Agents involved
- Director (Grok) led mapping, commits, scoring.
- Leonard = DS4FLASH0731 (first scored ingest).
Artifacts
docs/clients/overcome-fitness/data-inventory-v1.mddocs/clients/overcome-fitness/threat-register-v1.mddocs/clients/overcome-fitness/README.mdtools/gbp-snapshot-form.html(v1.1)docs/templates/gbp-snapshot-intake.md(v1.1)docs/agents/leonard-gbp-execution-card.md(v1.1)
Next for this client
- Human accept Draft Benchmark or request edits.
- Optional: re-run on RTX 3060 target model with same intake; compare scores.
- Optional: deeper social/website surfaces.
How to log a run
- After human review of an agent draft, add one row to the appropriate table.
- If the draft required material correction for an error (see definition), increment
errorsand setlast_error_date. - Do not count unscored exploratory chat as a run.
Last updated: 2026-08-03