Files
veripath/docs/operations/task-reliability-ledger.md
T

73 lines
4.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Task Reliability Ledger
**Purpose:** Make the ≥100 error-free executions rule and Leonard scoring **measurable**.
Until a task type has documented runs, human approval remains required before any client-facing publish for that task type.
**Error definition (default):** Output that invents facts/numbers, mis-assigns Evidence Tier, includes solution language in Layer 1/2, omits required Snapshot Date / recency, violates template structure, or would be unsafe to show a client without correction.
---
## Publish-gate counters (automation eligibility)
| task_type | runs | errors | last_error_date | notes |
|-----------|------|--------|----------------|-------|
| gbp_snapshot_ingest_r1_r3 | 2 | 0 | — | DS4FLASH0731 + Gemma4 both 9/10 PASS 2026-08-03 |
| data_inventory_draft_l1a | 1 | 0 | — | Overcome Fitness; human accepted Draft Benchmark 2026-08-03 |
| threat_register_draft_l2 | 1 | 0 | — | Overcome Fitness; human accepted Draft Benchmark 2026-08-03 |
| review_reply_draft | 0 | 0 | — | Not auto-publish |
| gbp_field_edit_draft | 0 | 0 | — | Not auto-publish |
| qa_answer_draft | 0 | 0 | — | Not auto-publish |
**Rule:** A task_type may bypass human approval **only after** ≥100 completed runs with **zero errors** for that exact task_type, documented here. Reliability does not transfer across task types.
---
## Leonard / agent scored runs (rubric 02 × 5 criteria)
| date | agent | client | task_type | score_total | zeros? | pass? | reviewer | notes |
|------|-------|--------|-----------|-------------|--------|-------|----------|-------|
| 2026-08-03 | DS4FLASH0731 (Leonard caveman) | Overcome Fitness | gbp_snapshot_ingest_r1_r3 | 9/10 | No | Yes | Director (Grok) | R1+R2 correct; R3 off. Phone redacted. Template 1/2. |
| 2026-08-03 | Gemma4 (Leonard caveman, RTX 3060 path) | Overcome Fitness | gbp_snapshot_ingest_r1_r3 | 9/10 | No | Yes | Director (Grok) | Same band as DS4FLASH. R1+R2 correct. Phone redacted. Template 1/2. |
**Pass:** no zeros and total ≥ 7/10.
---
## Process log (first real client)
### 2026-08-06 — Commit identity incident (dfd052d)
**What happened**
1. Commit `dfd052d` (AI Visibility Unified Framework v1, Artifact A/B templates, GLOSSARY additions) was authored and committed as `Ty <tybala@outlook.com>` by Leonard.
2. The repo's established agent convention is `Leonard <leonard@hermes.local>` (51 prior commits, author + committer, July 25 Aug 2). The commit should have used that identity.
3. Root cause: I matched the last 5 commits' visible author (all Ty-drafted that day) instead of the durable 51-commit convention. The identity was set explicitly (repo-local `git config`), so this was a deliberate wrong choice, not passive inheritance — the shared vpsadmin config is `Epictetus` and the sandboxed HOME has no identity at all.
4. Consequence: git blame on `dfd052d` reads as human-authored, erasing the human/agent distinction for that commit.
**Decision (Tony, 2026-08-06):** Leave `dfd052d` as-is — do NOT force-push/rewrite. The two-branch workflow was retired specifically to avoid git complexity; a force-push reintroduces it and launders the record. The honest move is to keep the misattributed commit in history and document the correction, which is what this entry does.
**Going-forward rule:** Every Leonard commit uses `Leonard <leonard@hermes.local>`, set per-repo, every time. Never commit agent work under a human identity — it destroys the git-blame audit trail that the "verify against git state, don't trust self-reports" principle depends on.
**Last updated:** 2026-08-06
---
### 2026-08-03 — Overcome Fitness (first real non-Phoenix cold audit)
**What happened**
1. Human filled GBP observations (Maps).
2. Intake corrected live; form/template/card → v1.1.
3. Data Inventory + Threat Register drafted (T-002 update, T-011 new).
4. Scored Leonard runs: DS4FLASH0731 9/10 PASS; Gemma4 9/10 PASS.
5. **Human accepted Draft Benchmark** (not Locked).
6. Next collection: social surfaces (Director-led).
**Last updated:** 2026-08-03
---
## How to log a run
1. After human review of an agent draft, add one row to the appropriate table.
2. If the draft required material correction for an error (see definition), increment `errors` and set `last_error_date`.
3. Do not count unscored exploratory chat as a run.