Files
veripath/docs/operations/task-reliability-ledger.md
T

3.8 KiB
Raw Blame History

Task Reliability Ledger

Purpose: Make the ≥100 error-free executions rule and Leonard scoring measurable.
Until a task type has documented runs, human approval remains required before any client-facing publish for that task type.

Error definition (default): Output that invents facts/numbers, mis-assigns Evidence Tier, includes solution language in Layer 1/2, omits required Snapshot Date / recency, violates template structure, or would be unsafe to show a client without correction.


Publish-gate counters (automation eligibility)

task_type runs errors last_error_date notes
gbp_snapshot_ingest_r1_r3 1 0 First scored run 2026-08-03 DS4FLASH0731; 9/10 PASS; minor phone redaction + template shape
data_inventory_draft_l1a 1 0 Overcome Fitness public baseline written 2026-08-03
threat_register_draft_l2 1 0 Overcome Fitness; T-002 updated + T-011 added via R1/R2
review_reply_draft 0 0 Not auto-publish
gbp_field_edit_draft 0 0 Not auto-publish
qa_answer_draft 0 0 Not auto-publish

Rule: A task_type may bypass human approval only after ≥100 completed runs with zero errors for that exact task_type, documented here. Reliability does not transfer across task types.


Leonard / agent scored runs (rubric 02 × 5 criteria)

date agent client task_type score_total zeros? pass? reviewer notes
2026-08-03 DS4FLASH0731 (Leonard caveman) Overcome Fitness gbp_snapshot_ingest_r1_r3 9/10 No Yes Director (Grok) R1+R2 correct; R3 correctly not fired. Phone partially redacted (not in intake). Flat table vs inventory Notes format (template 1/2). Severity “Standard” on completeness vs register Major.

Pass: no zeros and total ≥ 7/10.

Rubric detail (this run)
Evidence 2 · Tier 2 · Scope 2 · Recency 2 · Template 1


Process log (first real client)

2026-08-03 — Overcome Fitness (first real non-Phoenix cold audit)

What happened

  1. Human filled GBP observations (Maps) for Overcome Fitness.
  2. Intake corrected live (booking button is on website, not GBP; services = strength training only; Updates/Posts present but empty; claimed status Unknown).
  3. GBP snapshot form + intake template + Leonard execution card upgraded to v1.1 so those fields are captured next time without chat back-and-forth.
  4. Data Inventory GBP row written (Tier 2, Snapshot Date 2026-08-03).
  5. Threat Register updated: T-002 evidence refreshed (R1); T-011 added (R2 — unclear claimed + thin reviews).
  6. Client README updated. Status remains Draft Benchmark (not Locked).
  7. First scored Leonard run: DS4FLASH0731, inlined execution rules (no repo file access), caveman mode. 9/10 PASS. Logged above.

Agents involved

  • Director (Grok) led mapping, commits, scoring.
  • Leonard = DS4FLASH0731 (first scored ingest).

Artifacts

  • docs/clients/overcome-fitness/data-inventory-v1.md
  • docs/clients/overcome-fitness/threat-register-v1.md
  • docs/clients/overcome-fitness/README.md
  • tools/gbp-snapshot-form.html (v1.1)
  • docs/templates/gbp-snapshot-intake.md (v1.1)
  • docs/agents/leonard-gbp-execution-card.md (v1.1)

Next for this client

  • Human accept Draft Benchmark or request edits.
  • Optional: re-run on RTX 3060 target model with same intake; compare scores.
  • Optional: deeper social/website surfaces.

How to log a run

  1. After human review of an agent draft, add one row to the appropriate table.
  2. If the draft required material correction for an error (see definition), increment errors and set last_error_date.
  3. Do not count unscored exploratory chat as a run.

Last updated: 2026-08-03