# Athena-Oracle: Development Design Document **Version:** 0.1.0 **Date:** 2026-07-08 **Status:** Draft — pre-implementation **Source branch:** `MVP-milestone` --- ## 1. North Star > *Surfacing cross-source convergence and using falsification to distinguish real momentum from noise.* Athena is an autonomous research intelligence engine that ingests from multiple fragmented sources, detects when the same signals appear across independent channels, and uses decay-based falsification to separate genuine trends from one-day spikes. The system is model-agnostic, lightweight, and designed to run unattended on a resource-constrained VPS. ### How this design supports the north star | North Star Principle | Design Decision | Why | |---|---|---| | Cross-source convergence | Keyword co-occurrence matrix (Phase 1), embeddings + vector search (Phase 2) | Keyword co-occurrence is the simplest convergence detector: if the same entity appears in ≥3 independent sources within a time window, it's converging. Embeddings Phase 2 adds semantic convergence for signals that use different words but mean the same thing. | | Falsification over confirmation | Exponential decay scoring | A 7-day hard cutoff is blunt: some trends die in 48 hours, some take 60 days to validate. `score = base_score * e^(-λ * days_since_last_signal)` naturally scores dying trends low and sustained trends high without arbitrary day thresholds. | | Autonomous operation | Cron/scheduled pipeline + graceful degradation | The pipeline runs daily without human intervention. If Ollama is down, ingestion continues and summarization defers to the next run. If one adapter fails, the rest still run. | | Lightweight deployment | Python + SQLite + Flask + host-level Ollama | No PostgreSQL, no Elasticsearch, no Redis, no Kubernetes. A single Python process, a single SQLite file, and an external Ollama REST API. | ### MVP Success Criteria The MVP is considered successful when: 1. **Research Time**: Bob answers 'what's trending this week' in <15 min of research time total (across all sources). 2. **Signal Detection**: Over 7 days, the theme scan correctly identifies ≥1 real trend that has genuine cross-source convergence. 3. **Noise Rejection**: Over the same 7-day period, the falsification engine kills ≥1 false signal (a one-day spike with no sustained arrivals) to demonstrate decay-based filtering is working. --- ## 2. System Architecture ``` ┌─────────────────────────────────────────────────────┐ │ Athena-Oracle Pipeline (Python, ~150-500MB) │ │ │ │ ┌────────────┐ ┌─────────┐ ┌──────────────────┐ │ │ │ GitHub │ │ arXiv │ │ Reddit │ │ │ │ Adapter │ │ Adapter │ │ Adapter │ │ │ └──────┬─────┘ └────┬────┘ └────────┬─────────┘ │ │ │ │ │ │ │ ┌──────┴─────────────┴──────────────┴──────────┐ │ │ │ Deduplication (URL hash + title simhash) │ │ │ └────────────────────┬─────────────────────────┘ │ │ │ │ │ ┌────────────────────┴─────────────────────────┐ │ │ │ Theme Tagging │ │ │ │ Phase 1: Keyword co-occurrence + fixed seeds│ │ │ │ Phase 2: all-MiniLM embeddings + BERTopic │ │ │ └────────────────────┬─────────────────────────┘ │ │ │ │ │ ┌────────────────────┴─────────────────────────┐ │ │ │ Falsification Engine │ │ │ │ Exponential decay scoring per theme │ │ │ │ Convergence threshold: ≥3 independent sources│ │ │ └────────────────────┬─────────────────────────┘ │ │ │ │ │ ┌────────────────────┴─────────────────────────┐ │ │ │ SQLite │ │ │ │ entries table + run_log + FTS5 index │ │ │ │ Phase 2: + sqlite-vec extension │ │ │ └──────────────────────────────────────────────┘ │ │ │ │ Output layer: │ │ ┌──────────┐ ┌──────────────┐ ┌────────────────┐ │ │ │ Flask API│ │ File drops │ │ MCP server │ │ │ │ (Phase 4)│ │ (Phase 4) │ │ (Phase 6) │ │ │ └──────────┘ └──────────────┘ └────────────────┘ │ │ │ │ Structured JSON logging → stdout + file rotation │ │ Discord/Slack webhook on 2+ day adapter failure │ └─────────────────────────────────────────────────────┘ │ │ ┌───────────────────────────────┘ │ HTTP REST API ▼ ┌─────────────────────┐ │ Ollama (host-level) │ │ llama3.2:1b │ │ (summarization) │ └─────────────────────┘ ``` ### Memory budget | Component | Phase 1 | Phase 2 | |---|---|---| | Python runtime + deps | ~80 MB | ~80 MB | | SQLite (in-process) | ~10 MB | ~10 MB | | all-MiniLM embeddings | — | ~80 MB | | sqlite-vec extension | — | ~2 MB | | Flask | ~1 MB | ~1 MB | | **Pipeline total** | **~90 MB** | **~173 MB** | | Ollama + model (host-level) | ~2 GB | ~2 GB | | **System total** | **~2.1 GB** | **~2.2 GB** | The 150MB constraint applies to the pipeline process. The full system footprint including Ollama is ~2GB. ### Daily Run Flow 1. **Preflight** — Check disk space, verify DB integrity (`PRAGMA integrity_check`), load adapter config 2. **Ingest** — Run all 6 adapters in parallel, collect raw items per source 3. **Dedup** — Hash URLs, match against existing entries, insert only new items 4. **Theme Tag** — Run keyword co-occurrence against new items, tag themes 5. **Falsification** — Recompute decay scores for all active themes, kill dead theses 6. **Archive** — Move entries older than 90 days to archive table 7. **Report** — Write daily summary to `/output/`, update run_log ### Failure Modes | Component | Failure | Impact | Recovery | |---|---|---|---| | Adapter | Rate limit / 503 | Missing items from that source | Retry next cycle; pipeline continues | | Ollama | Down | Summaries skipped | Entries stored with `summary = null`, deferred to next run | | SQLite | Disk full | No writes | Alert via webhook; manual cleanup | | SQLite | Corruption | Data loss | Restore from last backup | | Network | Outbound blocked | All adapters fail | Alert; pipeline exits code 2 | --- ## 3. Data Layer ### 3.1 SQLite schema Core tables (from `schema.sql`): - **entries** — one row per ingested item, deduplicated by `(source, source_id)` - **run_log** — one row per pipeline run, tracks per-source success/failure - **FTS5 virtual table** — full-text search over `title` and `extracted_text` - **Phase 2: sqlite-vec** — vector index for semantic similarity queries ### 3.2 Why SQLite - Zero external dependency, single file, survives container restarts - FTS5 is built-in (no separate search engine) - Handles 100K-1M rows without performance issues - sqlite-vec extension adds vector search without a separate database - No connection pooling needed (single-writer pipeline) - Postgres/pgvector is premature optimization at this scale ### 3.3 Data retention - Raw entries: 90 days - Summaries and convergence scores: 365 days - Periodic `VACUUM` to reclaim space - `archive.py` handles cold storage rotation (deferred to Phase 3) ### 3.4 Backup Strategy - **Pre-schema-change backup**: Before any `ALTER TABLE` or schema modification: ```bash sqlite3 oracle.db '.backup oracle.db.bak' ``` Store `.bak` files with date suffix in `/backup/` (`oracle.db.bak-YYYYMMDD`). - **Daily compressed backup**: At 01:00 UTC (off-peak): ```bash tar -czf /backup/oracle.db.$(date +%Y%m%d).tar.gz oracle.db ``` Retain 30 days of backups; purge older: `find /backup -name '*.tar.gz' -mtime +30 -delete` - **Recovery**: Restore from backup with `cp /backup/oracle.db.bak-YYYYMMDD oracle.db`, verify with `PRAGMA integrity_check` ### 3.5 Schema Migration - Versioned migration files in `migrations/` directory (e.g., `001_initial_schema.sql`, `002_add_theme_tags.sql`) - Applied at startup: pipeline checks `schema_version` table, runs any unapplied migrations in order - Each migration is a single atomic SQL file; no partial migrations - Rollback: each migration includes a comment with the reverse SQL if needed --- ## 4. Adapter Layer ### 4.1 Source adapters (HTTP-only) | Adapter | API | Rate limit | Auth required | |---|---|---|---| | GitHub | REST API | 60/hr (unauth), 5000/hr (token) | `GITHUB_TOKEN` | | arXiv | REST API | 1 req/sec (polite) | No | | Reddit | RSS/JSON | ~100 req/min | No (but OAuth recommended) | | Hacker News | Firebase API | Unofficial, ~30 req/sec | No | | HuggingFace | REST API | Throttled if aggressive | `HUGGINGFACE_TOKEN` | | RSS Feeds | RSS XML | Varies | No | **Decision:** HTTP-only adapters, no Playwright/Selenium. All 6 sources have programmatic APIs. Playwright would add Chromium's 300MB+ overhead and fragility. ### 4.2 Deduplication arXiv papers appear on HN, Reddit, and Twitter. Without deduplication, the same signal is counted 3× and produces false convergence. - **Phase 1:** UNIQUE constraint on `(source, source_id)` + URL hash dedup across sources - **Phase 2:** SimHash/MinHash content fingerprinting for near-duplicate detection ### 4.3 Rate limiting and retries - Per-adapter rate limits enforced in the adapter class - `tenacity` library for exponential backoff on transient failures (429, 503, timeout) - One failing adapter does not kill the pipeline ### 4.4 Adapter Interface Contract All adapters must implement this minimal interface: - **name() → str**: Unique identifier for the source (e.g., `"arxiv"`) - **fetch(query, limit) → list[dict]**: Returns normalized items with required schema fields: - `source` (str): Source name matching `name()` - `source_id` (str): Unique per-source ID for deduplication - `title` (str): Human-readable title - `url` (str): Direct URL to the entry - `timestamp` (datetime): Publication/update time - `raw_score` (float): Source-specific signal strength (0.0–10.0) - `body` (str): Raw text content for summarization and theme tagging **Custom Exceptions:** - `RateLimitError`: Raised when source returns 429 or similar; includes `retry_after` (seconds) - `SourceUnavailableError`: Raised when source is down (5xx) or unreachable Adapters must not raise other exceptions on normal operation; unexpected errors should be logged with full traceback and surfaced via the run_log, not propagated to the pipeline. --- ## 5. Theme Tagging and Convergence Detection ### 5.1 Phase 1: Keyword co-occurrence Pre-defined keyword dictionaries per theme. An entry is tagged if ≥2 keywords from a theme dictionary appear in its title or extracted text. A theme "converges" if it appears in ≥3 independent sources within the last 24 hours. **Why keyword first at 150MB:** Keyword matching is zero-dependency, explainable, and works within the memory constraint. FTS5 provides fast retrieval. **Fixed themes:** The initial 4 themes (`tool-call`, `context`, `compute`, `trust`) are seeds, not a hard limit. An "other" catch-all bucket captures signals that don't match predefined themes. ### 5.2 Phase 2: Embeddings + auto-discovery - **all-MiniLM-L6-v2** (22M params, ~80MB) for sentence embeddings - **sqlite-vec** for in-database ANN search - **BERTopic** (or equivalent) for semi-supervised theme discovery, seeded from the Phase 1 dictionary - Hybrid query: FTS5 for precision (keyword match) + vector for recall (semantic match), merged via Reciprocal Rank Fusion **Why not keyword forever:** Keyword matching cannot detect semantic convergence (different words, same concept) and requires constant manual dictionary updates. Embeddings are the eventual target; Phase 1 is the bridge. ### 5.3 Convergence scoring ``` convergence_score = Σ(source_weights) × temporal_proximity × theme_entropy where: source_weights: arXiv=2.0, GitHub=1.5, HN=1.0, Reddit=0.8, HF=1.2, RSS=0.5 temporal_proximity: e^(-0.1 * hours_since_first_signal) theme_entropy: log2(number_of_independent_sources) ``` Thresholds: - `≥ 3.0` → "confirmed" trend - `≥ 1.5` → "emerging" signal - `< 1.5` → "noise" ### 5.4 Theme Governance Themes are defined in a single YAML file: `themes.yaml`. Each theme entry includes: - `name`: human-readable identifier (e.g., `tool-call`) - `keywords`: list of keyword patterns for Phase 1 regex matching - `owner`: responsible person (for quarterly review) - `created_at`: ISO date of creation - `status`: `active` or `deprecated` **Lifecycle:** - **Propose**: New themes require owner nomination + approval from project lead - **Review**: Active themes are reviewed quarterly; deprecated if <2 hits in 30 days - **Retire**: Deprecated themes are excluded from convergence scoring after 90 days - **Reinstate**: Deprecated themes can be reactivated if signals re-emerge The "other" catch-all bucket captures signals that don't match any active theme and is reviewed during quarterly theme audits for potential new theme creation. --- ## 6. Falsification Engine ### 6.1 Exponential decay scoring Replace the 7-day dead thesis rule with: ``` thesis_score = initial_score × e^(-λ × days_since_last_signal) where λ = 0.1 (configurable) ``` A thesis is "dead" when its score falls below a configurable threshold (default: 0.1), not when it hits a fixed day count. This naturally handles: - Fast-dying trends (score drops quickly) - Slow-burn trends (score stays elevated) - Revived trends (new signal resets the decay clock) ### 6.2 Cross-source validation A signal is flagged "unverified" if: - Only 1 source has primary (non-derivative) coverage - The signal appears only in echo chambers (e.g., HN upvotes ≠ real adoption) - A counter-narrative exists in the same time window ### 6.3 Calibration Process - **Initial parameters**: Start with λ = 0.1 (half-life ~7 days) for all themes at deployment - **Validation window**: After 30 days of live operation, validate against historical data - **Adjustment triggers**: - If >20% of confirmed real trends were falsely killed → decrease λ (e.g., to 0.05, slower decay) - If >30% of noise signals were incorrectly confirmed → increase λ (e.g., to 0.15, faster decay) - **Documentation**: Record calibration decisions in `calibration_log.md` with date, old/new λ values, and rationale Re-calibrate quarterly or whenever a major theme dictionary change is made. --- ## 7. Output Layer ### 7.1 Consumer interfaces (progressive rollout) | Consumer | Interface | Phase | |---|---|---| | **Bob** (trend tracker) | Flask REST API: `GET /trends?theme=&period=7d` | 4 | | **Alice** (content creator) | Daily file drops: `/output/YYYY-MM-DD/trends.yaml` | 4 | | **Sam** (Hermes agent) | MCP server: `oracle_search`, `oracle_trends`, `oracle_verdicts` | 6 | ### 7.2 REST API (Phase 4) Flask endpoints: - `GET /health` — pipeline status, last run time, adapter health - `GET /trends` — active themes with convergence scores - `GET /entries` — search entries (keyword + phase 2: semantic) - `GET /verdicts` — confirmed/dead theses - `GET /convergence` — cross-source convergence matrix ### 7.3 File drops (Phase 4) Daily structured output at a known path: ``` /output/YYYY-MM-DD/ trends.yaml # Human-readable daily digest signals.json # Structured machine-readable output verdicts.json # Confirmed/dead thesis list ``` ### 7.4 MCP server (Phase 6) MCP tools for Hermes agent integration: - `oracle_search(query, source, date_range)` — search entries - `oracle_trends(theme, convergence_threshold)` — get active trends - `oracle_verdicts(status)` — confirmed or dead theses - `oracle_latest(source)` — most recent entry per source ### 7.5 MCP Tool Signatures All MCP tools must implement these minimum signatures: - **get_trends()**: - Request: `{}` (no params) - Response: `{ "trends": [{"name": str, "score": float, "sources": [str], "decay_score": float}] }` - **search_entry(query: str, source: str | None = None)**: - Request: `{ "query": str, "source": str | null }` - Response: `{ "entries": [{"title": str, "url": str, "summary": str, "score": float}] }` - **get_convergence_report()**: - Request: `{}` - Response: `{ "converged": [{"entity": str, "sources": [str], "confidence": float}] }` Tools must validate input types and return empty arrays (not errors) for valid queries that yield no results. ### 7.6 API Authentication Model - **Phase 4 (MVP)**: Simple API key in `X-API-Key` header. No expiration, stored in config file (`/etc/athena/api_keys.yaml`) - **Phase 6 (Production)**: JWT bearer token with scopes (read-only, read-write, admin). Tokens expire after 24 hours; refresh via `/auth/token` endpoint Auth failures return HTTP 401 with `{ "error": "unauthorized" }`. Rate limiting applies per-key: 100 requests/minute. --- ## 8. Observability and Reliability ### 8.1 Logging Structured JSON logging via stdlib `logging` with JSON formatter. Per-pipeline-stage logs (ingest, dedup, theme, falsification, summarize) with source-level granularity. ### 8.2 Health endpoint `GET /health` returns: ```json { "status": "ok", "last_run": "2026-07-08T13:00:00Z", "last_run_duration_sec": 245, "entries_since_last_run": 127, "adapters": { "github": {"status": "ok", "fetched": 20}, "arxiv": {"status": "ok", "fetched": 15}, "reddit": {"status": "error", "fetched": 0, "error": "429 rate limited"} } } ``` ### 8.3 Alerting Discord/Slack webhook triggered when: - An adapter fails for 2+ consecutive days - Pipeline run exceeds 2× expected duration - SQLite database integrity check fails ### 8.4 Graceful degradation If Ollama is unreachable: - Ingestion continues normally - Summarization is skipped, entries stored with `summary = null` - **Deferral**: next run summarizes pending entries - **Alert**: "summarization deferred, N entries pending" ### 8.5 Log Retention and SLIs - **Log retention**: 30 days rolling; gzip-compressed after 7 days to save disk space - **SLI definitions**: - **Pipeline success rate**: >95% of daily runs complete without critical failure (exit code 2) - **Adapter availability**: >90% of scheduled runs successfully fetch each source (per-source metric) - **Theme detection accuracy**: ≥80% of manually verified trends identified correctly in first week ### 8.6 Alerting Matrix | Condition | Channel | Severity | Response Time | |---|---|---|---| | Adapter fails >2 consecutive runs | Discord webhook | P2 | Investigate within 1 hour | | Pipeline exit code 2 | Discord webhook + email | P1 | Investigate within 30 minutes | | DB disk usage >85% | Discord webhook | P2 | Investigate within 2 hours | | Ollama unreachable >5 min | Discord webhook | P2 | Restart service if needed | | Pipeline success rate <90% for 3 days | Email + dashboard | P3 | Review next cycle | Alerts are deduplicated: same condition won't fire again until resolved. --- ## 9. Scheduling ### 9.1 Phase 1: Cron - `oracle-pipeline.sh` invoked by cron at 13:00 UTC daily - `flock`/PID file prevents overlapping runs - Exit codes: 0 = success, 1 = partial failure, 2 = total failure ### 9.2 Phase 2: systemd timers - `Persistent=true` catches up on missed runs - `RandomizedDelaySec` prevents thundering herd - `OnFailureSec` for retry logic - Better logging than cron (`journalctl -u athena-timer`) ### 9.3 Why not APScheduler (Phase 1) APScheduler adds in-process async daemon overhead. Cron/systemd is OS-level, zero process memory cost, and sufficient for daily runs. APScheduler is the Phase 2 target if dynamic scheduling (user-configurable refresh rates, per-source intervals) is needed. ### 9.4 Scheduling Recommendation - **Phase 1 (MVP)**: Use system cron for daily runs at 13:00 UTC. Sufficient for fixed schedule, zero process memory cost, OS-level reliability. - **Phase 2**: Switch to systemd timers if dynamic scheduling needed (skip runs on holidays, adjust time zones). Better observability and integration with monitoring tools. - **APScheduler**: Only use if per-source intervals are required (e.g., arXiv every 6 hours, Reddit every 30 min). Adds ~50MB process overhead — not recommended for Phase 1. **Recommendation**: Start with cron for Phase 1. Evaluate whether Phase 2 sources need different intervals before considering systemd timers or APScheduler. --- ## 10. Inference ### 10.1 Summarization **Model:** Ollama `llama3.2:1b` (or `qwen2.5:0.5b` for lower resource) **Deployment:** Host-level Ollama service, pipeline calls via HTTP REST API **Contract:** Model-agnostic — `summarize(text) → (summary, model)` interface **Graceful degradation:** If Ollama is down, store raw text and defer summarization ### 10.2 Why not in-container Ollama Ollama daemon + 1B model requires ~2GB RAM. Running it inside the 150MB container is physically impossible. Running it host-level means the pipeline process stays within budget and Ollama can share resources with other services. ### 10.3 Model Upgrade Process - **Swap model**: Replace model file; update systemd unit `ExecStart` path if needed - **Restart service**: `systemctl restart qwythos-gpu0.service` (or equivalent unit) - **Verify health**: `curl http://localhost:8081/v1/health` — should return `{ "status": "ready" }` - **Quality check**: Run first summarization on known-good source (e.g., arXiv paper), compare output against baseline summary for content accuracy, length consistency, and hallucination rate - **Rollback**: If quality degrades (e.g., summary length <50 tokens, hallucination rate >10%), revert to previous model file immediately **Quality metrics**: Pass if summary is >50 tokens, no hallucinations on known entities, and theme detection matches golden samples. Only proceed with upgrade after successful verification. --- ## 11. Security | Requirement | Implementation | |---|---| | No hardcoded secrets | `GITHUB_TOKEN`, `HUGGINGFACE_TOKEN` as env vars or mounted secret files | | TLS for outbound | All HTTP adapters use `https://` | | Least privilege | Pipeline runs as standard user (no sudo) | | DB protection | `chmod 600 oracle.db` | | Input sanitization | Parameterized SQL queries, no string concatenation | ### 11.2 MVP Security Baseline > **Updated 2026-07-12:** Container-specific hardening below is *design intent*, not the current deployment. The live system runs as the `vpsadmin` standard user (no sudo) under Hermes cron. The equivalent protections that DO apply today: no hardcoded secrets (`.env` sourced by the pipeline scripts), DB file `chmod 600`, parameterized SQL. Container hardening is deferred to the future containerization task. - **Current (bare-script) baseline**: - Runs as `vpsadmin` standard user (no sudo) - Secrets in `/home/vpsadmin/oracle/.env` (`chmod 600`), never in code - DB protected: `chmod 600 oracle.db` - All outbound adapters use `https://` - **Deferred (container) hardening** *(design intent, not live)*: - Run as non-root user (`user: nobody` in Dockerfile) - Read-only filesystem where possible (except `/tmp`, `/var/log`) - No SSH access inside container; pipeline is cron-triggered, no interactive access needed - Minimal base image: `python:3.11-slim` (no dev tools, no git) - **Dependency scanning**: - CI pipeline runs `pip audit` or `safety check` on every push to `MVP-milestone` - Fail build if critical vulnerabilities found (>CVSS 7.0) - Warn on medium/high vulnerabilities; require manual review before merging - **Secret management**: - No secrets in code or config files (use environment variables at runtime) - API keys stored in `/etc/athena/secrets.yaml` with restricted permissions (`chmod 0600`) **Enforcement**: Security checks are automated in CI; local development is permissive but container builds must pass all scans. --- ## 12. Implementation Phases | Phase | Scope | Deliverable | |---|---|---| | **P0: Foundation** | schema.sql + sqlite-vec design, adapter registry, oracle-pipeline.sh skeleton | Empty but valid pipeline | | **P1: First data** | arXiv + RSS adapters, SQLite storage, keyword convergence, dedup | Live data flowing | | **P2: Full ingest** | GitHub, HN, HF, Reddit adapters, rate limiting, structured logging | All 6 sources live | | **P3: Falsification** | Exponential decay scoring, Ollama summarization, graceful degradation | Trend verdicts working | | **P4: Consumption** | Flask API, daily file drops, health endpoint, alerting webhooks | Bob and Alice can consume | | **P5: Validation** | 7-day UAT window, Hermes cron integration, exit codes | System runs unattended | | **P6: Scale** | sqlite-vec + embeddings, MCP server, APScheduler, BERTopic themes | Research-grade system | ### 12.4 Definition of Done per Phase Each phase must pass all listed criteria before being marked complete: **P0 (Foundation)**: - `schema.sql` creates all tables without errors - Empty pipeline runs cleanly with exit code 0 - `oracle-pipeline.sh` is idempotent (safe to run twice) - Adapter registry loads all 6 adapters **P1 (First data)**: - arXiv + RSS adapters fetch successfully - Entries stored in SQLite with correct schema - Keyword convergence detects at least 1 theme - Deduplication works (no duplicate entries) **P2 (Full ingest)**: - All 6 adapters fetch successfully in one run - Rate limiting enforced per adapter - Structured JSON logs emitted per pipeline stage - No adapter failure kills the pipeline **P3 (Falsification)**: - Exponential decay scoring implemented - Ollama summarization works with graceful degradation - Trend verdicts computed: confirmed/emerging/dead - Running over 7 days shows false signals dying **P4 (Consumption)**: - REST API endpoints return valid JSON - Daily file drops written to `/output/` - Health endpoint reports accurate status - Alerting webhooks fire on simulated failures **P5 (Validation)**: - 7-day UAT: pipeline runs unattended without intervention - Hermes cron integration works - Exit codes correct: 0=success, 1=partial, 2=failure - Pipeline completes <30 minutes end-to-end **P6 (Scale)**: - sqlite-vec + embeddings operational - MCP server responds to all 3 tools - APScheduler handles per-source intervals - System stays within 150MB pipeline memory budget **General**: No critical bugs open, all unit tests pass, CI green, security scan clean. --- ## 13. Backport to PRD: Requirements to Add The following requirements are implied by this design and should be added to `docs/MVP-PRD.md`: ### 13.1 Platform requirements (REQ-PLT-XX) | ID | Requirement | |---|---| | REQ-PLT-05 | All processes run as standard user (no sudo) — *already exists* | | REQ-PLT-10 | The pipeline process shall not exceed 500MB of RSS memory (excluding host-level Ollama) | | REQ-PLT-15 | The system shall support deployment on a VPS with 2GB total RAM (pipeline + Ollama + OS) | | REQ-PLT-20 | Ollama inference shall run as a host-level service, not inside the pipeline container | | REQ-PLT-25 | The pipeline shall use SQLite as the sole database (no PostgreSQL, no Elasticsearch, no Redis) | ### 13.2 Reliability requirements (REQ-REL-XX) | ID | Requirement | |---|---| | REQ-REL-05 | The system operates autonomously without human interaction — *already exists* | | REQ-REL-10 | Failed source fetches retry with exponential backoff — *already exists* | | REQ-REL-15 | Previously stored data is not lost on restart or failure — *already exists* | | REQ-REL-20 | The pipeline completes successfully even if 1+ sources are unavailable — *already exists* | | REQ-REL-25 | If Ollama is unreachable, ingestion continues and summarization defers to the next run | | REQ-REL-30 | The pipeline uses flock/PID file to prevent overlapping runs | ### 13.3 Observability requirements (REQ-DIAG-XX) | ID | Requirement | |---|---| | REQ-DIAG-05 | Structured JSON logs with timestamps and severity — *already exists* | | REQ-DIAG-10 | Health endpoint reports system status and last successful run — *already exists* | | REQ-DIAG-15 | Per-source success/failure and fetch counts logged per run — *already exists* | | REQ-DIAG-20 | Webhook alert (Discord/Slack) fires when an adapter fails for 2+ consecutive days | | REQ-DIAG-25 | run_log table captures per-run metrics queryable via SQL | ### 13.4 Integration requirements (REQ-INT-XX) | ID | Requirement | |---|---| | REQ-INT-05 | Structured, machine-readable output (JSON) consumable by external tools — *already exists* | | REQ-INT-10 | Adapter layer supports adding new sources without modifying core pipeline logic — *already exists* | | REQ-INT-15 | REST API exposes GET /trends, /entries, /verdicts endpoints | | REQ-INT-20 | Daily file drop at configurable path with structured output (YAML + JSON) | | REQ-INT-25 | MCP server exposes oracle_search, oracle_trends, oracle_verdicts tools (Phase 6) | ### 13.5 Data requirements (REQ-DATA-XX) *(new section)* | ID | Requirement | |---|---| | REQ-DATA-05 | Entries are deduplicated by (source, source_id) with cross-source URL hash deduplication | | REQ-DATA-10 | FTS5 full-text index on title and extracted_text for keyword search | | REQ-DATA-15 | Convergence detection: theme appears in ≥3 independent sources within 24h window | | REQ-DATA-20 | Falsification uses exponential decay scoring (configurable λ), not fixed day thresholds | | REQ-DATA-25 | Data retention: raw entries 90 days, summaries/convergence 365 days, periodic VACUUM | ### 13.6 Security requirements (REQ-SEC-XX) | ID | Requirement | |---|---| | REQ-SEC-05 | All outbound HTTP uses TLS — *already exists* | | REQ-SEC-10 | No secrets hardcoded or stored in plaintext — *already exists* | | REQ-SEC-15 | Least-privilege access for outbound API calls — *already exists* | | REQ-SEC-20 | SQLite database file permissions set to 600 (owner-only read/write) | | REQ-SEC-25 | All SQL queries use parameterized statements (no string concatenation) | ### 13.7 Scheduling requirements (REQ-SCH-XX) *(new section)* | ID | Requirement | |---|---| | REQ-SCH-05 | Pipeline entry point is oracle-pipeline.sh (idempotent, single command) | | REQ-SCH-10 | Default schedule: 13:00 UTC daily | | REQ-SCH-15 | Exit codes: 0 = success, 1 = partial failure, 2 = total failure | | REQ-SCH-20 | Overlapping run prevention via flock or PID file check | --- ## 14. Design Decisions Summary (Why) | Decision | Why | Revisit | |---|---|---| | **SQLite over PostgreSQL** | Zero external dependency, single file, FTS5 built-in, handles 1M rows fine. pgvector is premature at this scale. | P6 (if scale demands pgvector) | | **Ollama host-level** | 150MB container cannot fit Ollama + model (~2GB). Host-level lets pipeline stay within budget. | P2 (if inference needs change) | | **Flask over FastAPI** | ~1MB vs ~100MB runtime overhead. FastAPI is Phase 2 target; Flask suffices for internal REST API. | P2 (when FastAPI becomes viable) | | **Keyword co-occurrence (Phase 1)** | Zero-dependency, explainable, works at 150MB. Embeddings (Phase 2) add semantic convergence. | P2 (when embeddings ready) | | **Fixed themes + catch-all** | BERTopic requires 4GB RAM. Fixed themes with "other" bucket is the pragmatic constraint choice. | P6 (when auto-discovery needed) | | **Cron over APScheduler (Phase 1)** | OS-level, zero process memory cost. APScheduler is Phase 2 for dynamic scheduling. | P2 (if per-source intervals needed) | | **HTTP-only adapters** | All 6 sources have programmatic APIs. Playwright adds 300MB+ overhead and fragility. | N/A (stable) | | **Exponential decay over 7-day rule** | One-line formula, no fixed threshold. Handles fast-dying and slow-burn trends naturally. | N/A (stable) | | **Deduplication required** | arXiv papers appear on HN/Reddit/Twitter. Without dedup, same signal counted 3× = false convergence. | N/A (stable) | | **Graceful degradation on Ollama** | Ingestion must not depend on summarization. Store raw data, defer summaries. | N/A (stable) | ## 15. Testing Strategy ### 15.1 Unit Tests - **Per adapter**: Test each of the 6 adapters with known-good endpoints. Verify: - Returns valid JSON with required schema fields (source, source_id, title, url, timestamp) - Handles rate limits gracefully (no infinite loops) - Correct error codes for 429/503 - **Scoring functions**: Unit tests for exponential decay, convergence scoring, and theme matching. Include edge cases (empty input, negative scores). ### 15.2 Integration Tests - Run pipeline against seed dataset (100 entries from arXiv + RSS). Verify: - All adapters fetch successfully - Deduplication removes duplicates correctly - Theme detection identifies at least 3 themes - No critical errors in logs ### 15.3 E2E Tests - **Cron-to-snapshot**: Run full pipeline via cron, verify output files match expected snapshot (golden file comparison) - **Stress test**: Run 10 consecutive daily cycles with simulated failures (adapter timeout, Ollama down). Verify graceful degradation and recovery **Test coverage goal**: 80% of critical paths covered by automated tests. Manual testing for theme quality and summary accuracy. ## 16. Deployment and CI/CD ### 16.1 CI Pipeline > **Updated 2026-07-12:** CI originally built a Docker image per push. The live system has no CI/Docker build. The equivalent checks run manually/optionally: lint, `pip audit`, and `pytest` if tests exist. Container build is deferred. On every push to `MVP-milestone`: - **Lint**: `ruff check`, `mdlint docs/` - **Test**: Run unit tests (`pytest tests/`), integration tests with seed dataset - **Scan**: Run `pip audit`; fail if critical vulnerabilities (>CVSS 7.0) found - *(Deferred: Docker image build/tag with commit SHA)* ### 16.2 Deployment > **Deployment reality (updated 2026-07-12):** This document originally specified Docker containerization (150MB cap, non-root, `python:3.11-slim`). The deployed system does **not** use containers. It runs as bare Python scripts under the `vpsadmin` standard user, invoked by Hermes cron (`oracle-pipeline.sh` at 13:00 UTC). The 150MB budget, non-root, and read-only-filesystem hardening described below remain *design intent* and are tracked as a future containerization task — not the current state. Ollama is also not deployed; summarization is deferred (graceful degradation is the live behavior). - **Local dev / Production**: Bare Python on the VPS, no Docker. `pip install` deps, run `schema.sql`, source `.env`, then Hermes cron drives `oracle-pipeline.sh`. - **Ollama**: Not currently running. Pipeline stores raw text and defers summarization (see §10.2 / REQ-REL-25). Adding Ollama is a separate step, host-level. ### 16.3 Rollback Procedure If deployment fails or quality degrades: 1. Stop services: `systemctl stop oracle-pipeline qwythos-gpu0` 2. Restore previous code: `git checkout ` 3. Rebuild Docker image from restored code 4. Restart services: `systemctl start oracle-pipeline qwythos-gpu0` 5. Verify health: `curl http://localhost:8081/v1/health` **Rollback window**: Must complete within 5 minutes of failure detection. **Tagging**: Each deploy is tagged with semantic versioning (`v1.0.0`, `v1.1.0`) for easy rollback reference. ## 17. Operational Runbooks ### 17.1 Daily Pipeline Verification At 13:00 UTC after each run: 1. Check logs: `journalctl -u oracle-pipeline --since "today" | grep ERROR` 2. Verify exit code in `/var/log/oracle-pipeline/run_log.txt` — should be 0 3. Confirm output: `ls -lh /output/$(date +%Y-%m-%d)/` — should have `summary.json` and `metrics.json` 4. If any check fails, investigate with the relevant runbook below ### 17.2 Database Recovery from Corruption If `sqlite3 oracle.db 'PRAGMA integrity_check'` returns errors: 1. Stop pipeline: `systemctl stop oracle-pipeline` 2. Restore from last backup: `cp /backup/oracle.db.bak-YYYYMMDD oracle.db` 3. Verify integrity: `sqlite3 oracle.db 'PRAGMA integrity_check'` — should return "ok" 4. Restart pipeline: `systemctl start oracle-pipeline` **Prevention**: Daily compressed backups to `/backup/`, retention 30 days. ### 17.3 Adapter Failure Investigation If adapter fails repeatedly (>2 consecutive runs): 1. Check logs: `journalctl -u oracle-pipeline --since "today" | grep -A5 "Adapter"` 2. Test endpoint manually: `curl -X GET ` — verify HTTP status 3. Check rate limit headers: `curl -I | grep -i 'x-ratelimit'` 4. If rate-limited: Wait `retry_after` seconds, pipeline retries next cycle 5. Escalate to project lead if >3 consecutive failures ### 17.4 Ollama Service Restart If Ollama becomes unresponsive: 1. Check status: `systemctl status qwythos-gpu0.service` 2. View logs: `journalctl -u qwythos-gpu0.service --since "today"` 3. Restart service: `systemctl restart qwythos-gpu0.service` 4. Verify health: `curl http://localhost:8081/v1/health` — should return `{ "status": "ready" }` 5. Check GPU: `nvidia-smi` — ensure service actually loaded **Escalation**: If issue persists after restart, notify project lead. **Runbook maintenance**: Update runbooks when features change. Document changes in `runbook_changes.md`. ## 18. Risk Register and Assumptions ### 18.1 Risk Register | Risk | Likelihood | Impact | Mitigation | |---|---|---|---| | Source API changes (rate limits, endpoint shifts) | High | High | Auto-retry with exponential backoff; monitor adapter health; log API changes for review | | Model quality drift (summarization accuracy degrades) | Medium | Medium | Weekly golden sample comparison; rollback if hallucination rate >10% or summary <50 tokens | | Disk space exhaustion (backups, logs) | High | High | Automated 30-day retention; alert at 85% usage; compress old backups | | Single operator bottleneck (manual theme review) | Medium | Medium | Documented theme governance; quarterly review cycle; catch-all "other" bucket | | Network outage (all adapters fail) | Low | High | Graceful degradation: store raw data without summaries; resume when network returns | | SQLite corruption during schema migration | Low | Critical | Pre-migration backup; atomic migration scripts; integrity check after each migration | ### 18.2 Key Assumptions - All 6 source APIs remain stable for at least 6 months (no breaking changes) - GPU memory available for Qwythos inference (~15GB free on GPU0) - Network connectivity to all sources is available during pipeline runs - Single operator can complete manual theme review within 24 hours - Disk space sufficient for 30-day backup retention (~10GB) **Risk review**: Quarterly review of this register; update mitigations when new risks emerge. **Assumption tracking**: If an assumption proves false, document the deviation in `assumption_deviations.md` and reassess risks. ## 19. Dependencies and Tooling ### 19.1 Python Runtime Dependencies All dependencies pinned to exact versions in `requirements.txt`: - **requests**: HTTP client for all source APIs - **beautifulsoup4**: HTML parsing for RSS feeds - **feedparser**: RSS/Atom feed handling (primary) - **tenacity**: Retry logic with exponential backoff (used by all adapters) - **flask**: Internal REST API for Phase 4+ (served on port 8081) No external database dependencies — SQLite is built into Python. ### 19.2 Runtime System Dependencies - **Ollama**: Host-level inference service (~2GB RAM, GPU acceleration) - **systemd**: Service management for pipeline and inference (`oracle-pipeline.service`, `qwythos-gpu0.service`) - **cron**: Daily schedule trigger (Phase 1); systemd timers for Phase 2 - **sqlite3**: Database CLI for backup/recovery commands ### 19.3 Development Tooling - **ruff**: Linting and formatting (Python) - **pytest**: Unit and integration test framework - **docker**: Containerization for pipeline - **pip audit / safety**: Dependency vulnerability scanning (CI checks) ### 19.4 Dependency Update Policy - Critical security updates: Apply within 7 days of CVE disclosure - Minor version updates: Test in staging before production deployment - Major version upgrades: Require migration scripts and rollback plan **Tooling policy**: No new dependencies without approval from project lead; document rationale in `dependency_justifications.md`. **Lock file**: `requirements.txt` pinned to exact versions for reproducible builds. --- *Document prepared via multi-model analysis: Qwythos-9B (architectural critique), Qwen3.5-9B (implementation evaluation), and cross-review synthesis. 4 delegations, 2 rounds of debate. Raw reviews saved in the same directory.*