Files
veripath/docs/operations/task-reliability-ledger.md

4.2 KiB
Raw Permalink Blame History

Task Reliability Ledger

Purpose: Make the ≥100 error-free executions rule and Leonard scoring measurable.
Until a task type has documented runs, human approval remains required before any client-facing publish for that task type.

Error definition (default): Output that invents facts/numbers, mis-assigns Evidence Tier, includes solution language in Layer 1/2, omits required Snapshot Date / recency, violates template structure, or would be unsafe to show a client without correction.


Publish-gate counters (automation eligibility)

task_type runs errors last_error_date notes
gbp_snapshot_ingest_r1_r3 2 0 DS4FLASH0731 + Gemma4 both 9/10 PASS 2026-08-03
data_inventory_draft_l1a 1 0 Overcome Fitness; human accepted Draft Benchmark 2026-08-03
threat_register_draft_l2 1 0 Overcome Fitness; human accepted Draft Benchmark 2026-08-03
review_reply_draft 0 0 Not auto-publish
gbp_field_edit_draft 0 0 Not auto-publish
qa_answer_draft 0 0 Not auto-publish

Rule: A task_type may bypass human approval only after ≥100 completed runs with zero errors for that exact task_type, documented here. Reliability does not transfer across task types.


Leonard / agent scored runs (rubric 02 × 5 criteria)

date agent client task_type score_total zeros? pass? reviewer notes
2026-08-03 DS4FLASH0731 (Leonard caveman) Overcome Fitness gbp_snapshot_ingest_r1_r3 9/10 No Yes Director (Grok) R1+R2 correct; R3 off. Phone redacted. Template 1/2.
2026-08-03 Gemma4 (Leonard caveman, RTX 3060 path) Overcome Fitness gbp_snapshot_ingest_r1_r3 9/10 No Yes Director (Grok) Same band as DS4FLASH. R1+R2 correct. Phone redacted. Template 1/2.

Pass: no zeros and total ≥ 7/10.


Process log (first real client)

2026-08-06 — Commit identity incident (dfd052d)

What happened

  1. Commit dfd052d (AI Visibility Unified Framework v1, Artifact A/B templates, GLOSSARY additions) was authored and committed as Ty <tybala@outlook.com> by Leonard.
  2. The repo's established agent convention is Leonard <leonard@hermes.local> (51 prior commits, author + committer, July 25 Aug 2). The commit should have used that identity.
  3. Root cause: I matched the last 5 commits' visible author (all Ty-drafted that day) instead of the durable 51-commit convention. The identity was set explicitly (repo-local git config), so this was a deliberate wrong choice, not passive inheritance — the shared vpsadmin config is Epictetus and the sandboxed HOME has no identity at all.
  4. Consequence: git blame on dfd052d reads as human-authored, erasing the human/agent distinction for that commit.

Decision (Tony, 2026-08-06): Leave dfd052d as-is — do NOT force-push/rewrite. The two-branch workflow was retired specifically to avoid git complexity; a force-push reintroduces it and launders the record. The honest move is to keep the misattributed commit in history and document the correction, which is what this entry does.

Going-forward rule: Every Leonard commit uses Leonard <leonard@hermes.local>, set per-repo, every time. Never commit agent work under a human identity — it destroys the git-blame audit trail that the "verify against git state, don't trust self-reports" principle depends on.

Last updated: 2026-08-06


2026-08-03 — Overcome Fitness (first real non-Phoenix cold audit)

What happened

  1. Human filled GBP observations (Maps).
  2. Intake corrected live; form/template/card → v1.1.
  3. Data Inventory + Threat Register drafted (T-002 update, T-011 new).
  4. Scored Leonard runs: DS4FLASH0731 9/10 PASS; Gemma4 9/10 PASS.
  5. Human accepted Draft Benchmark (not Locked).
  6. Next collection: social surfaces (Director-led).

Last updated: 2026-08-03


How to log a run

  1. After human review of an agent draft, add one row to the appropriate table.
  2. If the draft required material correction for an error (see definition), increment errors and set last_error_date.
  3. Do not count unscored exploratory chat as a run.