# SPRINT 1 FINDINGS — Reference Material Raw evidence from the Sprint 1 manual review (200 stories, athena_review_report.txt). This document is **reference material for future decisions** — it records what the data actually showed so later tuning has a baseline. --- ## What We Learned Sprint 1 was pure rules (no embeddings, no LLM, no semantic similarity) by founder directive: *discover the taxonomy before adding intelligence*. The classifier's job was not accuracy — it was to surface misclassifications as signal. It did. ## High Signal (ship) Stories the editorial test would pass — reader can DO something: - **Local AI** — gguf / rtx / quantization / llama.cpp. Highest enthusiast + replication scores. - e.g. "I benched quad 5060Tis for code generation with Qwen3.6-27B" (final 0.50) - "Voodoo Quant beats Unsloth Dynamic 2.0 KLD by 95%" (0.45) - **Open-source tools** — things to install/run. - "Show HN: Juggler – open-source GUI coding agent" (0.57, highest in set) - "Open Source Local LLM Training Tool (consumer hardware)" (0.46) - "Zer0Fit — MCP server for zero-shot ML, 100% local" (0.30–0.46) - **Benchmarks with real numbers** — reproducible, comparable. - "GPUHedge: cold-start p95 117s→30s" (0.41) - "Migrating to GPT-5.6: 2.2x faster, 27% cheaper" (0.41) - **Infrastructure with code** — fixable, runnable. - "Your $80 Tesla P100 silently noisy math in llama.cpp — 3-line fix, free" (0.32) - **Production learnings** — transferable lessons. - "The real bottleneck for AI agents may be proving who they are" (0.19) - "AI agent crawlers now need permission" (0.25) ## Low Signal (reject) Stories the editorial test would fail — reader does NOTHING: - **Funding rounds** — PixVerse $439M, DeepSeek $7B. Money moved, no capability changed. - **Lawsuits** — Apple/OpenAI trade secret, Google training suit. Spectator content. - **CEO opinions** — Altman, Hassabis, Mosseri takes. Someone said something. - **Valuations** — PixVerse $2B. Number, no action. - **Corporate announcements** — Waze features, Spotify assistant, Siri beta. No reproducible substance. ## Taxonomy Gap (the real discovery) - **45 / 200 (22.5%) UNCATEGORIZED.** Mostly Reddit/HN opinion + culture-adjacent pieces with no keyword hit (e.g. "Stop Telling Me to Ask an LLM", "Mesh LLM", "Ghost Font"). - This is **not a failure** — it's the editorial boundary made visible. The keyword set is narrower than taste. Tuning Sprint 2 (Lens) must start from these 45, not from classifier precision. ## Bucket distribution (200 reviewed) | Bucket | n | |--------|---| | RESEARCH | 40 | | INFRASTRUCTURE | 23 | | UNCATEGORIZED | 45 | | MODEL RELEASE | 24 | | PROBLEM SOLVED | 19 | | SHIPPING | 12 | | CULTURE | 15 | | BUSINESS | 13 | | LOCAL AI | 9 | ## Open question for Sprint 2 When the Lens arrives, its first job is **not** to classify better. It's to explain *why* a story is or isn't actionable — and to propose a section (§SECTION_DEFINITIONS) for human confirmation. The 45 UNCATEGORIZED are the training set for that judgment, not the 155 that already matched.