diff --git a/docs/operations/task-reliability-ledger.md b/docs/operations/task-reliability-ledger.md index 4aa19a6..2fa5e7a 100644 --- a/docs/operations/task-reliability-ledger.md +++ b/docs/operations/task-reliability-ledger.md @@ -11,9 +11,9 @@ Until a task type has documented runs, human approval remains required before an | task_type | runs | errors | last_error_date | notes | |-----------|------|--------|----------------|-------| -| gbp_snapshot_ingest_r1_r3 | 1 | 0 | — | First scored run 2026-08-03 DS4FLASH0731; 9/10 PASS; minor phone redaction + template shape | -| data_inventory_draft_l1a | 1 | 0 | — | Overcome Fitness public baseline written 2026-08-03 | -| threat_register_draft_l2 | 1 | 0 | — | Overcome Fitness; T-002 updated + T-011 added via R1/R2 | +| gbp_snapshot_ingest_r1_r3 | 2 | 0 | — | DS4FLASH0731 + Gemma4 both 9/10 PASS 2026-08-03 | +| data_inventory_draft_l1a | 1 | 0 | — | Overcome Fitness; human accepted Draft Benchmark 2026-08-03 | +| threat_register_draft_l2 | 1 | 0 | — | Overcome Fitness; human accepted Draft Benchmark 2026-08-03 | | review_reply_draft | 0 | 0 | — | Not auto-publish | | gbp_field_edit_draft | 0 | 0 | — | Not auto-publish | | qa_answer_draft | 0 | 0 | — | Not auto-publish | @@ -26,13 +26,11 @@ Until a task type has documented runs, human approval remains required before an | date | agent | client | task_type | score_total | zeros? | pass? | reviewer | notes | |------|-------|--------|-----------|-------------|--------|-------|----------|-------| -| 2026-08-03 | DS4FLASH0731 (Leonard caveman) | Overcome Fitness | gbp_snapshot_ingest_r1_r3 | 9/10 | No | Yes | Director (Grok) | R1+R2 correct; R3 correctly not fired. Phone partially redacted (not in intake). Flat table vs inventory Notes format (template 1/2). Severity “Standard” on completeness vs register Major. | +| 2026-08-03 | DS4FLASH0731 (Leonard caveman) | Overcome Fitness | gbp_snapshot_ingest_r1_r3 | 9/10 | No | Yes | Director (Grok) | R1+R2 correct; R3 off. Phone redacted. Template 1/2. | +| 2026-08-03 | Gemma4 (Leonard caveman, RTX 3060 path) | Overcome Fitness | gbp_snapshot_ingest_r1_r3 | 9/10 | No | Yes | Director (Grok) | Same band as DS4FLASH. R1+R2 correct. Phone redacted. Template 1/2. | **Pass:** no zeros and total ≥ 7/10. -**Rubric detail (this run)** -Evidence 2 · Tier 2 · Scope 2 · Recency 2 · Template 1 - --- ## Process log (first real client) @@ -40,30 +38,14 @@ Evidence 2 · Tier 2 · Scope 2 · Recency 2 · Template 1 ### 2026-08-03 — Overcome Fitness (first real non-Phoenix cold audit) **What happened** -1. Human filled GBP observations (Maps) for Overcome Fitness. -2. Intake corrected live (booking button is on website, not GBP; services = strength training only; Updates/Posts present but empty; claimed status Unknown). -3. GBP snapshot form + intake template + Leonard execution card upgraded to **v1.1** so those fields are captured next time without chat back-and-forth. -4. Data Inventory GBP row written (Tier 2, Snapshot Date 2026-08-03). -5. Threat Register updated: T-002 evidence refreshed (R1); **T-011** added (R2 — unclear claimed + thin reviews). -6. Client README updated. Status remains **Draft Benchmark** (not Locked). -7. **First scored Leonard run:** DS4FLASH0731, inlined execution rules (no repo file access), caveman mode. **9/10 PASS.** Logged above. +1. Human filled GBP observations (Maps). +2. Intake corrected live; form/template/card → v1.1. +3. Data Inventory + Threat Register drafted (T-002 update, T-011 new). +4. Scored Leonard runs: DS4FLASH0731 9/10 PASS; Gemma4 9/10 PASS. +5. **Human accepted Draft Benchmark** (not Locked). +6. Next collection: social surfaces (Director-led). -**Agents involved** -- Director (Grok) led mapping, commits, scoring. -- Leonard = DS4FLASH0731 (first scored ingest). - -**Artifacts** -- `docs/clients/overcome-fitness/data-inventory-v1.md` -- `docs/clients/overcome-fitness/threat-register-v1.md` -- `docs/clients/overcome-fitness/README.md` -- `tools/gbp-snapshot-form.html` (v1.1) -- `docs/templates/gbp-snapshot-intake.md` (v1.1) -- `docs/agents/leonard-gbp-execution-card.md` (v1.1) - -**Next for this client** -- Human accept Draft Benchmark or request edits. -- Optional: re-run on RTX 3060 target model with same intake; compare scores. -- Optional: deeper social/website surfaces. +**Last updated:** 2026-08-03 --- @@ -72,5 +54,3 @@ Evidence 2 · Tier 2 · Scope 2 · Recency 2 · Template 1 1. After human review of an agent draft, add one row to the appropriate table. 2. If the draft required material correction for an error (see definition), increment `errors` and set `last_error_date`. 3. Do not count unscored exploratory chat as a run. - -**Last updated:** 2026-08-03