Multi-model architecture review (Qwythos-9B + Qwen3.5-9B, 2 rounds): - System architecture with memory budget (150MB pipeline + 2GB system) - SQLite + FTS5 data layer with Phase 2 sqlite-vec upgrade path - Keyword co-occurrence convergence (P1) → embeddings (P2) - Exponential decay falsification replacing 7-day fixed rule - Progressive output layer: file drops → Flask API → MCP server - 6-phase implementation plan (P0-P6) with deliverables - 30+ PRD backport requirements organized by category - Design decisions rationale table
22 KiB
Athena-Oracle: Development Design Document
Version: 0.1.0
Date: 2026-07-08
Status: Draft — pre-implementation
Source branch: MVP-milestone
1. North Star
Surfacing cross-source convergence and using falsification to distinguish real momentum from noise.
Athena is an autonomous research intelligence engine that ingests from multiple fragmented sources, detects when the same signals appear across independent channels, and uses decay-based falsification to separate genuine trends from one-day spikes. The system is model-agnostic, lightweight, and designed to run unattended on a resource-constrained VPS.
How this design supports the north star
| North Star Principle | Design Decision | Why |
|---|---|---|
| Cross-source convergence | Keyword co-occurrence matrix (Phase 1), embeddings + vector search (Phase 2) | Keyword co-occurrence is the simplest convergence detector: if the same entity appears in ≥3 independent sources within a time window, it's converging. Embeddings Phase 2 adds semantic convergence for signals that use different words but mean the same thing. |
| Falsification over confirmation | Exponential decay scoring | A 7-day hard cutoff is blunt: some trends die in 48 hours, some take 60 days to validate. score = base_score * e^(-λ * days_since_last_signal) naturally scores dying trends low and sustained trends high without arbitrary day thresholds. |
| Autonomous operation | Cron/scheduled pipeline + graceful degradation | The pipeline runs daily without human intervention. If Ollama is down, ingestion continues and summarization defers to the next run. If one adapter fails, the rest still run. |
| Lightweight deployment | Python + SQLite + Flask + host-level Ollama | No PostgreSQL, no Elasticsearch, no Redis, no Kubernetes. A single Python process, a single SQLite file, and an external Ollama REST API. |
2. System Architecture
┌─────────────────────────────────────────────────────┐
│ Athena-Oracle Pipeline (Python, ~150-500MB) │
│ │
│ ┌────────────┐ ┌─────────┐ ┌──────────────────┐ │
│ │ GitHub │ │ arXiv │ │ Reddit │ │
│ │ Adapter │ │ Adapter │ │ Adapter │ │
│ └──────┬─────┘ └────┬────┘ └────────┬─────────┘ │
│ │ │ │ │
│ ┌──────┴─────────────┴──────────────┴──────────┐ │
│ │ Deduplication (URL hash + title simhash) │ │
│ └────────────────────┬─────────────────────────┘ │
│ │ │
│ ┌────────────────────┴─────────────────────────┐ │
│ │ Theme Tagging │ │
│ │ Phase 1: Keyword co-occurrence + fixed seeds│ │
│ │ Phase 2: all-MiniLM embeddings + BERTopic │ │
│ └────────────────────┬─────────────────────────┘ │
│ │ │
│ ┌────────────────────┴─────────────────────────┐ │
│ │ Falsification Engine │ │
│ │ Exponential decay scoring per theme │ │
│ │ Convergence threshold: ≥3 independent sources│ │
│ └────────────────────┬─────────────────────────┘ │
│ │ │
│ ┌────────────────────┴─────────────────────────┐ │
│ │ SQLite │ │
│ │ entries table + run_log + FTS5 index │ │
│ │ Phase 2: + sqlite-vec extension │ │
│ └──────────────────────────────────────────────┘ │
│ │
│ Output layer: │
│ ┌──────────┐ ┌──────────────┐ ┌────────────────┐ │
│ │ Flask API│ │ File drops │ │ MCP server │ │
│ │ (Phase 4)│ │ (Phase 4) │ │ (Phase 6) │ │
│ └──────────┘ └──────────────┘ └────────────────┘ │
│ │
│ Structured JSON logging → stdout + file rotation │
│ Discord/Slack webhook on 2+ day adapter failure │
└─────────────────────────────────────────────────────┘
│
│
┌───────────────────────────────┘
│ HTTP REST API
▼
┌─────────────────────┐
│ Ollama (host-level) │
│ llama3.2:1b │
│ (summarization) │
└─────────────────────┘
Memory budget
| Component | Phase 1 | Phase 2 |
|---|---|---|
| Python runtime + deps | ~80 MB | ~80 MB |
| SQLite (in-process) | ~10 MB | ~10 MB |
| all-MiniLM embeddings | — | ~80 MB |
| sqlite-vec extension | — | ~2 MB |
| Flask | ~1 MB | ~1 MB |
| Pipeline total | ~90 MB | ~173 MB |
| Ollama + model (host-level) | ~2 GB | ~2 GB |
| System total | ~2.1 GB | ~2.2 GB |
The 150MB constraint applies to the pipeline process. The full system footprint including Ollama is ~2GB.
3. Data Layer
3.1 SQLite schema
Core tables (from schema.sql):
- entries — one row per ingested item, deduplicated by
(source, source_id) - run_log — one row per pipeline run, tracks per-source success/failure
- FTS5 virtual table — full-text search over
titleandextracted_text - Phase 2: sqlite-vec — vector index for semantic similarity queries
3.2 Why SQLite
- Zero external dependency, single file, survives container restarts
- FTS5 is built-in (no separate search engine)
- Handles 100K-1M rows without performance issues
- sqlite-vec extension adds vector search without a separate database
- No connection pooling needed (single-writer pipeline)
- Postgres/pgvector is premature optimization at this scale
3.3 Data retention
- Raw entries: 90 days
- Summaries and convergence scores: 365 days
- Periodic
VACUUMto reclaim space archive.pyhandles cold storage rotation (deferred to Phase 3)
4. Adapter Layer
4.1 Source adapters (HTTP-only)
| Adapter | API | Rate limit | Auth required |
|---|---|---|---|
| GitHub | REST API | 60/hr (unauth), 5000/hr (token) | GITHUB_TOKEN |
| arXiv | REST API | 1 req/sec (polite) | No |
| RSS/JSON | ~100 req/min | No (but OAuth recommended) | |
| Hacker News | Firebase API | Unofficial, ~30 req/sec | No |
| HuggingFace | REST API | Throttled if aggressive | HUGGINGFACE_TOKEN |
| RSS Feeds | RSS XML | Varies | No |
Decision: HTTP-only adapters, no Playwright/Selenium. All 6 sources have programmatic APIs. Playwright would add Chromium's 300MB+ overhead and fragility.
4.2 Deduplication
arXiv papers appear on HN, Reddit, and Twitter. Without deduplication, the same signal is counted 3× and produces false convergence.
- Phase 1: UNIQUE constraint on
(source, source_id)+ URL hash dedup across sources - Phase 2: SimHash/MinHash content fingerprinting for near-duplicate detection
4.3 Rate limiting and retries
- Per-adapter rate limits enforced in the adapter class
tenacitylibrary for exponential backoff on transient failures (429, 503, timeout)- One failing adapter does not kill the pipeline
5. Theme Tagging and Convergence Detection
5.1 Phase 1: Keyword co-occurrence
Pre-defined keyword dictionaries per theme. An entry is tagged if ≥2 keywords from a theme dictionary appear in its title or extracted text. A theme "converges" if it appears in ≥3 independent sources within the last 24 hours.
Why keyword first at 150MB: Keyword matching is zero-dependency, explainable, and works within the memory constraint. FTS5 provides fast retrieval.
Fixed themes: The initial 4 themes (tool-call, context, compute, trust) are seeds, not a hard limit. An "other" catch-all bucket captures signals that don't match predefined themes.
5.2 Phase 2: Embeddings + auto-discovery
- all-MiniLM-L6-v2 (22M params, ~80MB) for sentence embeddings
- sqlite-vec for in-database ANN search
- BERTopic (or equivalent) for semi-supervised theme discovery, seeded from the Phase 1 dictionary
- Hybrid query: FTS5 for precision (keyword match) + vector for recall (semantic match), merged via Reciprocal Rank Fusion
Why not keyword forever: Keyword matching cannot detect semantic convergence (different words, same concept) and requires constant manual dictionary updates. Embeddings are the eventual target; Phase 1 is the bridge.
5.3 Convergence scoring
convergence_score = Σ(source_weights) × temporal_proximity × theme_entropy
where:
source_weights: arXiv=2.0, GitHub=1.5, HN=1.0, Reddit=0.8, HF=1.2, RSS=0.5
temporal_proximity: e^(-0.1 * hours_since_first_signal)
theme_entropy: log2(number_of_independent_sources)
Thresholds:
≥ 3.0→ "confirmed" trend≥ 1.5→ "emerging" signal< 1.5→ "noise"
6. Falsification Engine
6.1 Exponential decay scoring
Replace the 7-day dead thesis rule with:
thesis_score = initial_score × e^(-λ × days_since_last_signal)
where λ = 0.1 (configurable)
A thesis is "dead" when its score falls below a configurable threshold (default: 0.1), not when it hits a fixed day count. This naturally handles:
- Fast-dying trends (score drops quickly)
- Slow-burn trends (score stays elevated)
- Revived trends (new signal resets the decay clock)
6.2 Cross-source validation
A signal is flagged "unverified" if:
- Only 1 source has primary (non-derivative) coverage
- The signal appears only in echo chambers (e.g., HN upvotes ≠ real adoption)
- A counter-narrative exists in the same time window
7. Output Layer
7.1 Consumer interfaces (progressive rollout)
| Consumer | Interface | Phase |
|---|---|---|
| Bob (trend tracker) | Flask REST API: GET /trends?theme=&period=7d |
4 |
| Alice (content creator) | Daily file drops: /output/YYYY-MM-DD/trends.yaml |
4 |
| Sam (Hermes agent) | MCP server: oracle_search, oracle_trends, oracle_verdicts |
6 |
7.2 REST API (Phase 4)
Flask endpoints:
GET /health— pipeline status, last run time, adapter healthGET /trends— active themes with convergence scoresGET /entries— search entries (keyword + phase 2: semantic)GET /verdicts— confirmed/dead thesesGET /convergence— cross-source convergence matrix
7.3 File drops (Phase 4)
Daily structured output at a known path:
/output/YYYY-MM-DD/
trends.yaml # Human-readable daily digest
signals.json # Structured machine-readable output
verdicts.json # Confirmed/dead thesis list
7.4 MCP server (Phase 6)
MCP tools for Hermes agent integration:
oracle_search(query, source, date_range)— search entriesoracle_trends(theme, convergence_threshold)— get active trendsoracle_verdicts(status)— confirmed or dead thesesoracle_latest(source)— most recent entry per source
8. Observability and Reliability
8.1 Logging
Structured JSON logging via stdlib logging with JSON formatter. Per-pipeline-stage logs (ingest, dedup, theme, falsification, summarize) with source-level granularity.
8.2 Health endpoint
GET /health returns:
{
"status": "ok",
"last_run": "2026-07-08T13:00:00Z",
"last_run_duration_sec": 245,
"entries_since_last_run": 127,
"adapters": {
"github": {"status": "ok", "fetched": 20},
"arxiv": {"status": "ok", "fetched": 15},
"reddit": {"status": "error", "fetched": 0, "error": "429 rate limited"}
}
}
8.3 Alerting
Discord/Slack webhook triggered when:
- An adapter fails for 2+ consecutive days
- Pipeline run exceeds 2× expected duration
- SQLite database integrity check fails
8.4 Graceful degradation
If Ollama is unreachable:
- Ingestion continues normally
- Summarization is skipped, entries stored with
summary = null - Deferral: next run summarizes pending entries
- Alert: "summarization deferred, N entries pending"
9. Scheduling
9.1 Phase 1: Cron
oracle-pipeline.shinvoked by cron at 13:00 UTC dailyflock/PID file prevents overlapping runs- Exit codes: 0 = success, 1 = partial failure, 2 = total failure
9.2 Phase 2: systemd timers
Persistent=truecatches up on missed runsRandomizedDelaySecprevents thundering herdOnFailureSecfor retry logic- Better logging than cron (
journalctl -u athena-timer)
9.3 Why not APScheduler (Phase 1)
APScheduler adds in-process async daemon overhead. Cron/systemd is OS-level, zero process memory cost, and sufficient for daily runs. APScheduler is the Phase 2 target if dynamic scheduling (user-configurable refresh rates, per-source intervals) is needed.
10. Inference
10.1 Summarization
Model: Ollama llama3.2:1b (or qwen2.5:0.5b for lower resource)
Deployment: Host-level Ollama service, pipeline calls via HTTP REST API
Contract: Model-agnostic — summarize(text) → (summary, model) interface
Graceful degradation: If Ollama is down, store raw text and defer summarization
10.2 Why not in-container Ollama
Ollama daemon + 1B model requires ~2GB RAM. Running it inside the 150MB container is physically impossible. Running it host-level means the pipeline process stays within budget and Ollama can share resources with other services.
11. Security
| Requirement | Implementation |
|---|---|
| No hardcoded secrets | GITHUB_TOKEN, HUGGINGFACE_TOKEN as env vars or mounted secret files |
| TLS for outbound | All HTTP adapters use https:// |
| Least privilege | Pipeline runs as standard user (no sudo) |
| DB protection | chmod 600 oracle.db |
| Input sanitization | Parameterized SQL queries, no string concatenation |
12. Implementation Phases
| Phase | Scope | Deliverable |
|---|---|---|
| P0: Foundation | schema.sql + sqlite-vec design, adapter registry, oracle-pipeline.sh skeleton | Empty but valid pipeline |
| P1: First data | arXiv + RSS adapters, SQLite storage, keyword convergence, dedup | Live data flowing |
| P2: Full ingest | GitHub, HN, HF, Reddit adapters, rate limiting, structured logging | All 6 sources live |
| P3: Falsification | Exponential decay scoring, Ollama summarization, graceful degradation | Trend verdicts working |
| P4: Consumption | Flask API, daily file drops, health endpoint, alerting webhooks | Bob and Alice can consume |
| P5: Validation | 7-day UAT window, Hermes cron integration, exit codes | System runs unattended |
| P6: Scale | sqlite-vec + embeddings, MCP server, APScheduler, BERTopic themes | Research-grade system |
13. Backport to PRD: Requirements to Add
The following requirements are implied by this design and should be added to docs/MVP-PRD.md:
13.1 Platform requirements (REQ-PLT-XX)
| ID | Requirement |
|---|---|
| REQ-PLT-05 | All processes run as standard user (no sudo) — already exists |
| REQ-PLT-10 | The pipeline process shall not exceed 500MB of RSS memory (excluding host-level Ollama) |
| REQ-PLT-15 | The system shall support deployment on a VPS with 2GB total RAM (pipeline + Ollama + OS) |
| REQ-PLT-20 | Ollama inference shall run as a host-level service, not inside the pipeline container |
| REQ-PLT-25 | The pipeline shall use SQLite as the sole database (no PostgreSQL, no Elasticsearch, no Redis) |
13.2 Reliability requirements (REQ-REL-XX)
| ID | Requirement |
|---|---|
| REQ-REL-05 | The system operates autonomously without human interaction — already exists |
| REQ-REL-10 | Failed source fetches retry with exponential backoff — already exists |
| REQ-REL-15 | Previously stored data is not lost on restart or failure — already exists |
| REQ-REL-20 | The pipeline completes successfully even if 1+ sources are unavailable — already exists |
| REQ-REL-25 | If Ollama is unreachable, ingestion continues and summarization defers to the next run |
| REQ-REL-30 | The pipeline uses flock/PID file to prevent overlapping runs |
13.3 Observability requirements (REQ-DIAG-XX)
| ID | Requirement |
|---|---|
| REQ-DIAG-05 | Structured JSON logs with timestamps and severity — already exists |
| REQ-DIAG-10 | Health endpoint reports system status and last successful run — already exists |
| REQ-DIAG-15 | Per-source success/failure and fetch counts logged per run — already exists |
| REQ-DIAG-20 | Webhook alert (Discord/Slack) fires when an adapter fails for 2+ consecutive days |
| REQ-DIAG-25 | run_log table captures per-run metrics queryable via SQL |
13.4 Integration requirements (REQ-INT-XX)
| ID | Requirement |
|---|---|
| REQ-INT-05 | Structured, machine-readable output (JSON) consumable by external tools — already exists |
| REQ-INT-10 | Adapter layer supports adding new sources without modifying core pipeline logic — already exists |
| REQ-INT-15 | REST API exposes GET /trends, /entries, /verdicts endpoints |
| REQ-INT-20 | Daily file drop at configurable path with structured output (YAML + JSON) |
| REQ-INT-25 | MCP server exposes oracle_search, oracle_trends, oracle_verdicts tools (Phase 6) |
13.5 Data requirements (REQ-DATA-XX) (new section)
| ID | Requirement |
|---|---|
| REQ-DATA-05 | Entries are deduplicated by (source, source_id) with cross-source URL hash deduplication |
| REQ-DATA-10 | FTS5 full-text index on title and extracted_text for keyword search |
| REQ-DATA-15 | Convergence detection: theme appears in ≥3 independent sources within 24h window |
| REQ-DATA-20 | Falsification uses exponential decay scoring (configurable λ), not fixed day thresholds |
| REQ-DATA-25 | Data retention: raw entries 90 days, summaries/convergence 365 days, periodic VACUUM |
13.6 Security requirements (REQ-SEC-XX)
| ID | Requirement |
|---|---|
| REQ-SEC-05 | All outbound HTTP uses TLS — already exists |
| REQ-SEC-10 | No secrets hardcoded or stored in plaintext — already exists |
| REQ-SEC-15 | Least-privilege access for outbound API calls — already exists |
| REQ-SEC-20 | SQLite database file permissions set to 600 (owner-only read/write) |
| REQ-SEC-25 | All SQL queries use parameterized statements (no string concatenation) |
13.7 Scheduling requirements (REQ-SCH-XX) (new section)
| ID | Requirement |
|---|---|
| REQ-SCH-05 | Pipeline entry point is oracle-pipeline.sh (idempotent, single command) |
| REQ-SCH-10 | Default schedule: 13:00 UTC daily |
| REQ-SCH-15 | Exit codes: 0 = success, 1 = partial failure, 2 = total failure |
| REQ-SCH-20 | Overlapping run prevention via flock or PID file check |
14. Design Decisions Summary (Why)
| Decision | Why |
|---|---|
| SQLite over PostgreSQL | Zero external dependency, single file, FTS5 built-in, handles 1M rows fine. pgvector is premature at this scale. |
| Ollama host-level | 150MB container cannot fit Ollama + model (~2GB). Host-level lets pipeline stay within budget. |
| Flask over FastAPI | ~1MB vs ~100MB runtime overhead. FastAPI is Phase 2 target; Flask suffices for internal REST API. |
| Keyword co-occurrence (Phase 1) | Zero-dependency, explainable, works at 150MB. Embeddings (Phase 2) add semantic convergence. |
| Fixed themes + catch-all | BERTopic requires 4GB RAM. Fixed themes with "other" bucket is the pragmatic constraint choice. |
| Cron over APScheduler (Phase 1) | OS-level, zero process memory cost. APScheduler is Phase 2 for dynamic scheduling. |
| HTTP-only adapters | All 6 sources have programmatic APIs. Playwright adds 300MB+ overhead and fragility. |
| Exponential decay over 7-day rule | One-line formula, no fixed threshold. Handles fast-dying and slow-burn trends naturally. |
| Deduplication required | arXiv papers appear on HN/Reddit/Twitter. Without dedup, same signal counted 3× = false convergence. |
| Graceful degradation on Ollama | Ingestion must not depend on summarization. Store raw data, defer summaries. |
Document prepared via multi-model analysis: Qwythos-9B (architectural critique), Qwen3.5-9B (implementation evaluation), and cross-review synthesis. 4 delegations, 2 rounds of debate. Raw reviews saved in the same directory.