2.7 KiB
2.7 KiB
Task Reliability Ledger
Purpose: Make the ≥100 error-free executions rule and Leonard scoring measurable.
Until a task type has documented runs, human approval remains required before any client-facing publish for that task type.
Error definition (default): Output that invents facts/numbers, mis-assigns Evidence Tier, includes solution language in Layer 1/2, omits required Snapshot Date / recency, violates template structure, or would be unsafe to show a client without correction.
Publish-gate counters (automation eligibility)
| task_type | runs | errors | last_error_date | notes |
|---|---|---|---|---|
| gbp_snapshot_ingest_r1_r3 | 2 | 0 | — | DS4FLASH0731 + Gemma4 both 9/10 PASS 2026-08-03 |
| data_inventory_draft_l1a | 1 | 0 | — | Overcome Fitness; human accepted Draft Benchmark 2026-08-03 |
| threat_register_draft_l2 | 1 | 0 | — | Overcome Fitness; human accepted Draft Benchmark 2026-08-03 |
| review_reply_draft | 0 | 0 | — | Not auto-publish |
| gbp_field_edit_draft | 0 | 0 | — | Not auto-publish |
| qa_answer_draft | 0 | 0 | — | Not auto-publish |
Rule: A task_type may bypass human approval only after ≥100 completed runs with zero errors for that exact task_type, documented here. Reliability does not transfer across task types.
Leonard / agent scored runs (rubric 0–2 × 5 criteria)
| date | agent | client | task_type | score_total | zeros? | pass? | reviewer | notes |
|---|---|---|---|---|---|---|---|---|
| 2026-08-03 | DS4FLASH0731 (Leonard caveman) | Overcome Fitness | gbp_snapshot_ingest_r1_r3 | 9/10 | No | Yes | Director (Grok) | R1+R2 correct; R3 off. Phone redacted. Template 1/2. |
| 2026-08-03 | Gemma4 (Leonard caveman, RTX 3060 path) | Overcome Fitness | gbp_snapshot_ingest_r1_r3 | 9/10 | No | Yes | Director (Grok) | Same band as DS4FLASH. R1+R2 correct. Phone redacted. Template 1/2. |
Pass: no zeros and total ≥ 7/10.
Process log (first real client)
2026-08-03 — Overcome Fitness (first real non-Phoenix cold audit)
What happened
- Human filled GBP observations (Maps).
- Intake corrected live; form/template/card → v1.1.
- Data Inventory + Threat Register drafted (T-002 update, T-011 new).
- Scored Leonard runs: DS4FLASH0731 9/10 PASS; Gemma4 9/10 PASS.
- Human accepted Draft Benchmark (not Locked).
- Next collection: social surfaces (Director-led).
Last updated: 2026-08-03
How to log a run
- After human review of an agent draft, add one row to the appropriate table.
- If the draft required material correction for an error (see definition), increment
errorsand setlast_error_date. - Do not count unscored exploratory chat as a run.