Decision 2026-07-12: do NOT containerize. Document actual deployment state: - Bare Python scripts under vpsadmin user, Hermes cron drives oracle-pipeline.sh - Container hardening / 150MB cap / Docker build retained as future target (design intent) - Ollama not deployed; summarization deferred (graceful degradation is live) - .env (chmod 600) is the live secret mechanism, not /etc/athena/secrets.yaml
42 KiB
Athena-Oracle: Development Design Document
Version: 0.1.0
Date: 2026-07-08
Status: Draft — pre-implementation
Source branch: MVP-milestone
1. North Star
Surfacing cross-source convergence and using falsification to distinguish real momentum from noise.
Athena is an autonomous research intelligence engine that ingests from multiple fragmented sources, detects when the same signals appear across independent channels, and uses decay-based falsification to separate genuine trends from one-day spikes. The system is model-agnostic, lightweight, and designed to run unattended on a resource-constrained VPS.
How this design supports the north star
| North Star Principle | Design Decision | Why |
|---|---|---|
| Cross-source convergence | Keyword co-occurrence matrix (Phase 1), embeddings + vector search (Phase 2) | Keyword co-occurrence is the simplest convergence detector: if the same entity appears in ≥3 independent sources within a time window, it's converging. Embeddings Phase 2 adds semantic convergence for signals that use different words but mean the same thing. |
| Falsification over confirmation | Exponential decay scoring | A 7-day hard cutoff is blunt: some trends die in 48 hours, some take 60 days to validate. score = base_score * e^(-λ * days_since_last_signal) naturally scores dying trends low and sustained trends high without arbitrary day thresholds. |
| Autonomous operation | Cron/scheduled pipeline + graceful degradation | The pipeline runs daily without human intervention. If Ollama is down, ingestion continues and summarization defers to the next run. If one adapter fails, the rest still run. |
| Lightweight deployment | Python + SQLite + Flask + host-level Ollama | No PostgreSQL, no Elasticsearch, no Redis, no Kubernetes. A single Python process, a single SQLite file, and an external Ollama REST API. |
MVP Success Criteria
The MVP is considered successful when:
-
Research Time: Bob answers 'what's trending this week' in <15 min of research time total (across all sources).
-
Signal Detection: Over 7 days, the theme scan correctly identifies ≥1 real trend that has genuine cross-source convergence.
-
Noise Rejection: Over the same 7-day period, the falsification engine kills ≥1 false signal (a one-day spike with no sustained arrivals) to demonstrate decay-based filtering is working.
2. System Architecture
┌─────────────────────────────────────────────────────┐
│ Athena-Oracle Pipeline (Python, ~150-500MB) │
│ │
│ ┌────────────┐ ┌─────────┐ ┌──────────────────┐ │
│ │ GitHub │ │ arXiv │ │ Reddit │ │
│ │ Adapter │ │ Adapter │ │ Adapter │ │
│ └──────┬─────┘ └────┬────┘ └────────┬─────────┘ │
│ │ │ │ │
│ ┌──────┴─────────────┴──────────────┴──────────┐ │
│ │ Deduplication (URL hash + title simhash) │ │
│ └────────────────────┬─────────────────────────┘ │
│ │ │
│ ┌────────────────────┴─────────────────────────┐ │
│ │ Theme Tagging │ │
│ │ Phase 1: Keyword co-occurrence + fixed seeds│ │
│ │ Phase 2: all-MiniLM embeddings + BERTopic │ │
│ └────────────────────┬─────────────────────────┘ │
│ │ │
│ ┌────────────────────┴─────────────────────────┐ │
│ │ Falsification Engine │ │
│ │ Exponential decay scoring per theme │ │
│ │ Convergence threshold: ≥3 independent sources│ │
│ └────────────────────┬─────────────────────────┘ │
│ │ │
│ ┌────────────────────┴─────────────────────────┐ │
│ │ SQLite │ │
│ │ entries table + run_log + FTS5 index │ │
│ │ Phase 2: + sqlite-vec extension │ │
│ └──────────────────────────────────────────────┘ │
│ │
│ Output layer: │
│ ┌──────────┐ ┌──────────────┐ ┌────────────────┐ │
│ │ Flask API│ │ File drops │ │ MCP server │ │
│ │ (Phase 4)│ │ (Phase 4) │ │ (Phase 6) │ │
│ └──────────┘ └──────────────┘ └────────────────┘ │
│ │
│ Structured JSON logging → stdout + file rotation │
│ Discord/Slack webhook on 2+ day adapter failure │
└─────────────────────────────────────────────────────┘
│
│
┌───────────────────────────────┘
│ HTTP REST API
▼
┌─────────────────────┐
│ Ollama (host-level) │
│ llama3.2:1b │
│ (summarization) │
└─────────────────────┘
Memory budget
| Component | Phase 1 | Phase 2 |
|---|---|---|
| Python runtime + deps | ~80 MB | ~80 MB |
| SQLite (in-process) | ~10 MB | ~10 MB |
| all-MiniLM embeddings | — | ~80 MB |
| sqlite-vec extension | — | ~2 MB |
| Flask | ~1 MB | ~1 MB |
| Pipeline total | ~90 MB | ~173 MB |
| Ollama + model (host-level) | ~2 GB | ~2 GB |
| System total | ~2.1 GB | ~2.2 GB |
The 150MB constraint applies to the pipeline process. The full system footprint including Ollama is ~2GB.
Daily Run Flow
- Preflight — Check disk space, verify DB integrity (
PRAGMA integrity_check), load adapter config - Ingest — Run all 6 adapters in parallel, collect raw items per source
- Dedup — Hash URLs, match against existing entries, insert only new items
- Theme Tag — Run keyword co-occurrence against new items, tag themes
- Falsification — Recompute decay scores for all active themes, kill dead theses
- Archive — Move entries older than 90 days to archive table
- Report — Write daily summary to
/output/, update run_log
Failure Modes
| Component | Failure | Impact | Recovery |
|---|---|---|---|
| Adapter | Rate limit / 503 | Missing items from that source | Retry next cycle; pipeline continues |
| Ollama | Down | Summaries skipped | Entries stored with summary = null, deferred to next run |
| SQLite | Disk full | No writes | Alert via webhook; manual cleanup |
| SQLite | Corruption | Data loss | Restore from last backup |
| Network | Outbound blocked | All adapters fail | Alert; pipeline exits code 2 |
3. Data Layer
3.1 SQLite schema
Core tables (from schema.sql):
- entries — one row per ingested item, deduplicated by
(source, source_id) - run_log — one row per pipeline run, tracks per-source success/failure
- FTS5 virtual table — full-text search over
titleandextracted_text - Phase 2: sqlite-vec — vector index for semantic similarity queries
3.2 Why SQLite
- Zero external dependency, single file, survives container restarts
- FTS5 is built-in (no separate search engine)
- Handles 100K-1M rows without performance issues
- sqlite-vec extension adds vector search without a separate database
- No connection pooling needed (single-writer pipeline)
- Postgres/pgvector is premature optimization at this scale
3.3 Data retention
- Raw entries: 90 days
- Summaries and convergence scores: 365 days
- Periodic
VACUUMto reclaim space archive.pyhandles cold storage rotation (deferred to Phase 3)
3.4 Backup Strategy
-
Pre-schema-change backup: Before any
ALTER TABLEor schema modification:sqlite3 oracle.db '.backup oracle.db.bak'Store
.bakfiles with date suffix in/backup/(oracle.db.bak-YYYYMMDD). -
Daily compressed backup: At 01:00 UTC (off-peak):
tar -czf /backup/oracle.db.$(date +%Y%m%d).tar.gz oracle.dbRetain 30 days of backups; purge older:
find /backup -name '*.tar.gz' -mtime +30 -delete -
Recovery: Restore from backup with
cp /backup/oracle.db.bak-YYYYMMDD oracle.db, verify withPRAGMA integrity_check
3.5 Schema Migration
- Versioned migration files in
migrations/directory (e.g.,001_initial_schema.sql,002_add_theme_tags.sql) - Applied at startup: pipeline checks
schema_versiontable, runs any unapplied migrations in order - Each migration is a single atomic SQL file; no partial migrations
- Rollback: each migration includes a comment with the reverse SQL if needed
4. Adapter Layer
4.1 Source adapters (HTTP-only)
| Adapter | API | Rate limit | Auth required |
|---|---|---|---|
| GitHub | REST API | 60/hr (unauth), 5000/hr (token) | GITHUB_TOKEN |
| arXiv | REST API | 1 req/sec (polite) | No |
| RSS/JSON | ~100 req/min | No (but OAuth recommended) | |
| Hacker News | Firebase API | Unofficial, ~30 req/sec | No |
| HuggingFace | REST API | Throttled if aggressive | HUGGINGFACE_TOKEN |
| RSS Feeds | RSS XML | Varies | No |
Decision: HTTP-only adapters, no Playwright/Selenium. All 6 sources have programmatic APIs. Playwright would add Chromium's 300MB+ overhead and fragility.
4.2 Deduplication
arXiv papers appear on HN, Reddit, and Twitter. Without deduplication, the same signal is counted 3× and produces false convergence.
- Phase 1: UNIQUE constraint on
(source, source_id)+ URL hash dedup across sources - Phase 2: SimHash/MinHash content fingerprinting for near-duplicate detection
4.3 Rate limiting and retries
- Per-adapter rate limits enforced in the adapter class
tenacitylibrary for exponential backoff on transient failures (429, 503, timeout)- One failing adapter does not kill the pipeline
4.4 Adapter Interface Contract
All adapters must implement this minimal interface:
- name() → str: Unique identifier for the source (e.g.,
"arxiv") - fetch(query, limit) → list[dict]: Returns normalized items with required schema fields:
source(str): Source name matchingname()source_id(str): Unique per-source ID for deduplicationtitle(str): Human-readable titleurl(str): Direct URL to the entrytimestamp(datetime): Publication/update timeraw_score(float): Source-specific signal strength (0.0–10.0)body(str): Raw text content for summarization and theme tagging
Custom Exceptions:
RateLimitError: Raised when source returns 429 or similar; includesretry_after(seconds)SourceUnavailableError: Raised when source is down (5xx) or unreachable
Adapters must not raise other exceptions on normal operation; unexpected errors should be logged with full traceback and surfaced via the run_log, not propagated to the pipeline.
5. Theme Tagging and Convergence Detection
5.1 Phase 1: Keyword co-occurrence
Pre-defined keyword dictionaries per theme. An entry is tagged if ≥2 keywords from a theme dictionary appear in its title or extracted text. A theme "converges" if it appears in ≥3 independent sources within the last 24 hours.
Why keyword first at 150MB: Keyword matching is zero-dependency, explainable, and works within the memory constraint. FTS5 provides fast retrieval.
Fixed themes: The initial 4 themes (tool-call, context, compute, trust) are seeds, not a hard limit. An "other" catch-all bucket captures signals that don't match predefined themes.
5.2 Phase 2: Embeddings + auto-discovery
- all-MiniLM-L6-v2 (22M params, ~80MB) for sentence embeddings
- sqlite-vec for in-database ANN search
- BERTopic (or equivalent) for semi-supervised theme discovery, seeded from the Phase 1 dictionary
- Hybrid query: FTS5 for precision (keyword match) + vector for recall (semantic match), merged via Reciprocal Rank Fusion
Why not keyword forever: Keyword matching cannot detect semantic convergence (different words, same concept) and requires constant manual dictionary updates. Embeddings are the eventual target; Phase 1 is the bridge.
5.3 Convergence scoring
convergence_score = Σ(source_weights) × temporal_proximity × theme_entropy
where:
source_weights: arXiv=2.0, GitHub=1.5, HN=1.0, Reddit=0.8, HF=1.2, RSS=0.5
temporal_proximity: e^(-0.1 * hours_since_first_signal)
theme_entropy: log2(number_of_independent_sources)
Thresholds:
≥ 3.0→ "confirmed" trend≥ 1.5→ "emerging" signal< 1.5→ "noise"
5.4 Theme Governance
Themes are defined in a single YAML file: themes.yaml. Each theme entry includes:
name: human-readable identifier (e.g.,tool-call)keywords: list of keyword patterns for Phase 1 regex matchingowner: responsible person (for quarterly review)created_at: ISO date of creationstatus:activeordeprecated
Lifecycle:
- Propose: New themes require owner nomination + approval from project lead
- Review: Active themes are reviewed quarterly; deprecated if <2 hits in 30 days
- Retire: Deprecated themes are excluded from convergence scoring after 90 days
- Reinstate: Deprecated themes can be reactivated if signals re-emerge
The "other" catch-all bucket captures signals that don't match any active theme and is reviewed during quarterly theme audits for potential new theme creation.
6. Falsification Engine
6.1 Exponential decay scoring
Replace the 7-day dead thesis rule with:
thesis_score = initial_score × e^(-λ × days_since_last_signal)
where λ = 0.1 (configurable)
A thesis is "dead" when its score falls below a configurable threshold (default: 0.1), not when it hits a fixed day count. This naturally handles:
- Fast-dying trends (score drops quickly)
- Slow-burn trends (score stays elevated)
- Revived trends (new signal resets the decay clock)
6.2 Cross-source validation
A signal is flagged "unverified" if:
- Only 1 source has primary (non-derivative) coverage
- The signal appears only in echo chambers (e.g., HN upvotes ≠ real adoption)
- A counter-narrative exists in the same time window
6.3 Calibration Process
- Initial parameters: Start with λ = 0.1 (half-life ~7 days) for all themes at deployment
- Validation window: After 30 days of live operation, validate against historical data
- Adjustment triggers:
- If >20% of confirmed real trends were falsely killed → decrease λ (e.g., to 0.05, slower decay)
- If >30% of noise signals were incorrectly confirmed → increase λ (e.g., to 0.15, faster decay)
- Documentation: Record calibration decisions in
calibration_log.mdwith date, old/new λ values, and rationale
Re-calibrate quarterly or whenever a major theme dictionary change is made.
7. Output Layer
7.1 Consumer interfaces (progressive rollout)
| Consumer | Interface | Phase |
|---|---|---|
| Bob (trend tracker) | Flask REST API: GET /trends?theme=&period=7d |
4 |
| Alice (content creator) | Daily file drops: /output/YYYY-MM-DD/trends.yaml |
4 |
| Sam (Hermes agent) | MCP server: oracle_search, oracle_trends, oracle_verdicts |
6 |
7.2 REST API (Phase 4)
Flask endpoints:
GET /health— pipeline status, last run time, adapter healthGET /trends— active themes with convergence scoresGET /entries— search entries (keyword + phase 2: semantic)GET /verdicts— confirmed/dead thesesGET /convergence— cross-source convergence matrix
7.3 File drops (Phase 4)
Daily structured output at a known path:
/output/YYYY-MM-DD/
trends.yaml # Human-readable daily digest
signals.json # Structured machine-readable output
verdicts.json # Confirmed/dead thesis list
7.4 MCP server (Phase 6)
MCP tools for Hermes agent integration:
oracle_search(query, source, date_range)— search entriesoracle_trends(theme, convergence_threshold)— get active trendsoracle_verdicts(status)— confirmed or dead thesesoracle_latest(source)— most recent entry per source
7.5 MCP Tool Signatures
All MCP tools must implement these minimum signatures:
-
get_trends():
- Request:
{}(no params) - Response:
{ "trends": [{"name": str, "score": float, "sources": [str], "decay_score": float}] }
- Request:
-
search_entry(query: str, source: str | None = None):
- Request:
{ "query": str, "source": str | null } - Response:
{ "entries": [{"title": str, "url": str, "summary": str, "score": float}] }
- Request:
-
get_convergence_report():
- Request:
{} - Response:
{ "converged": [{"entity": str, "sources": [str], "confidence": float}] }
- Request:
Tools must validate input types and return empty arrays (not errors) for valid queries that yield no results.
7.6 API Authentication Model
- Phase 4 (MVP): Simple API key in
X-API-Keyheader. No expiration, stored in config file (/etc/athena/api_keys.yaml) - Phase 6 (Production): JWT bearer token with scopes (read-only, read-write, admin). Tokens expire after 24 hours; refresh via
/auth/tokenendpoint
Auth failures return HTTP 401 with { "error": "unauthorized" }. Rate limiting applies per-key: 100 requests/minute.
8. Observability and Reliability
8.1 Logging
Structured JSON logging via stdlib logging with JSON formatter. Per-pipeline-stage logs (ingest, dedup, theme, falsification, summarize) with source-level granularity.
8.2 Health endpoint
GET /health returns:
{
"status": "ok",
"last_run": "2026-07-08T13:00:00Z",
"last_run_duration_sec": 245,
"entries_since_last_run": 127,
"adapters": {
"github": {"status": "ok", "fetched": 20},
"arxiv": {"status": "ok", "fetched": 15},
"reddit": {"status": "error", "fetched": 0, "error": "429 rate limited"}
}
}
8.3 Alerting
Discord/Slack webhook triggered when:
- An adapter fails for 2+ consecutive days
- Pipeline run exceeds 2× expected duration
- SQLite database integrity check fails
8.4 Graceful degradation
If Ollama is unreachable:
- Ingestion continues normally
- Summarization is skipped, entries stored with
summary = null - Deferral: next run summarizes pending entries
- Alert: "summarization deferred, N entries pending"
8.5 Log Retention and SLIs
- Log retention: 30 days rolling; gzip-compressed after 7 days to save disk space
- SLI definitions:
- Pipeline success rate: >95% of daily runs complete without critical failure (exit code 2)
- Adapter availability: >90% of scheduled runs successfully fetch each source (per-source metric)
- Theme detection accuracy: ≥80% of manually verified trends identified correctly in first week
8.6 Alerting Matrix
| Condition | Channel | Severity | Response Time |
|---|---|---|---|
| Adapter fails >2 consecutive runs | Discord webhook | P2 | Investigate within 1 hour |
| Pipeline exit code 2 | Discord webhook + email | P1 | Investigate within 30 minutes |
| DB disk usage >85% | Discord webhook | P2 | Investigate within 2 hours |
| Ollama unreachable >5 min | Discord webhook | P2 | Restart service if needed |
| Pipeline success rate <90% for 3 days | Email + dashboard | P3 | Review next cycle |
Alerts are deduplicated: same condition won't fire again until resolved.
9. Scheduling
9.1 Phase 1: Cron
oracle-pipeline.shinvoked by cron at 13:00 UTC dailyflock/PID file prevents overlapping runs- Exit codes: 0 = success, 1 = partial failure, 2 = total failure
9.2 Phase 2: systemd timers
Persistent=truecatches up on missed runsRandomizedDelaySecprevents thundering herdOnFailureSecfor retry logic- Better logging than cron (
journalctl -u athena-timer)
9.3 Why not APScheduler (Phase 1)
APScheduler adds in-process async daemon overhead. Cron/systemd is OS-level, zero process memory cost, and sufficient for daily runs. APScheduler is the Phase 2 target if dynamic scheduling (user-configurable refresh rates, per-source intervals) is needed.
9.4 Scheduling Recommendation
-
Phase 1 (MVP): Use system cron for daily runs at 13:00 UTC. Sufficient for fixed schedule, zero process memory cost, OS-level reliability.
-
Phase 2: Switch to systemd timers if dynamic scheduling needed (skip runs on holidays, adjust time zones). Better observability and integration with monitoring tools.
-
APScheduler: Only use if per-source intervals are required (e.g., arXiv every 6 hours, Reddit every 30 min). Adds ~50MB process overhead — not recommended for Phase 1.
Recommendation: Start with cron for Phase 1. Evaluate whether Phase 2 sources need different intervals before considering systemd timers or APScheduler.
10. Inference
10.1 Summarization
Model: Ollama llama3.2:1b (or qwen2.5:0.5b for lower resource)
Deployment: Host-level Ollama service, pipeline calls via HTTP REST API
Contract: Model-agnostic — summarize(text) → (summary, model) interface
Graceful degradation: If Ollama is down, store raw text and defer summarization
10.2 Why not in-container Ollama
Ollama daemon + 1B model requires ~2GB RAM. Running it inside the 150MB container is physically impossible. Running it host-level means the pipeline process stays within budget and Ollama can share resources with other services.
10.3 Model Upgrade Process
- Swap model: Replace model file; update systemd unit
ExecStartpath if needed - Restart service:
systemctl restart qwythos-gpu0.service(or equivalent unit) - Verify health:
curl http://localhost:8081/v1/health— should return{ "status": "ready" } - Quality check: Run first summarization on known-good source (e.g., arXiv paper), compare output against baseline summary for content accuracy, length consistency, and hallucination rate
- Rollback: If quality degrades (e.g., summary length <50 tokens, hallucination rate >10%), revert to previous model file immediately
Quality metrics: Pass if summary is >50 tokens, no hallucinations on known entities, and theme detection matches golden samples. Only proceed with upgrade after successful verification.
11. Security
| Requirement | Implementation |
|---|---|
| No hardcoded secrets | GITHUB_TOKEN, HUGGINGFACE_TOKEN as env vars or mounted secret files |
| TLS for outbound | All HTTP adapters use https:// |
| Least privilege | Pipeline runs as standard user (no sudo) |
| DB protection | chmod 600 oracle.db |
| Input sanitization | Parameterized SQL queries, no string concatenation |
11.2 MVP Security Baseline
Updated 2026-07-12: Container-specific hardening below is design intent, not the current deployment. The live system runs as the
vpsadminstandard user (no sudo) under Hermes cron. The equivalent protections that DO apply today: no hardcoded secrets (.envsourced by the pipeline scripts), DB filechmod 600, parameterized SQL. Container hardening is deferred to the future containerization task.
-
Current (bare-script) baseline:
- Runs as
vpsadminstandard user (no sudo) - Secrets in
/home/vpsadmin/oracle/.env(chmod 600), never in code - DB protected:
chmod 600 oracle.db - All outbound adapters use
https://
- Runs as
-
Deferred (container) hardening (design intent, not live):
- Run as non-root user (
user: nobodyin Dockerfile) - Read-only filesystem where possible (except
/tmp,/var/log) - No SSH access inside container; pipeline is cron-triggered, no interactive access needed
- Minimal base image:
python:3.11-slim(no dev tools, no git)
- Run as non-root user (
-
Dependency scanning:
- CI pipeline runs
pip auditorsafety checkon every push toMVP-milestone - Fail build if critical vulnerabilities found (>CVSS 7.0)
- Warn on medium/high vulnerabilities; require manual review before merging
- CI pipeline runs
-
Secret management:
- No secrets in code or config files (use environment variables at runtime)
- API keys stored in
/etc/athena/secrets.yamlwith restricted permissions (chmod 0600)
Enforcement: Security checks are automated in CI; local development is permissive but container builds must pass all scans.
12. Implementation Phases
| Phase | Scope | Deliverable |
|---|---|---|
| P0: Foundation | schema.sql + sqlite-vec design, adapter registry, oracle-pipeline.sh skeleton | Empty but valid pipeline |
| P1: First data | arXiv + RSS adapters, SQLite storage, keyword convergence, dedup | Live data flowing |
| P2: Full ingest | GitHub, HN, HF, Reddit adapters, rate limiting, structured logging | All 6 sources live |
| P3: Falsification | Exponential decay scoring, Ollama summarization, graceful degradation | Trend verdicts working |
| P4: Consumption | Flask API, daily file drops, health endpoint, alerting webhooks | Bob and Alice can consume |
| P5: Validation | 7-day UAT window, Hermes cron integration, exit codes | System runs unattended |
| P6: Scale | sqlite-vec + embeddings, MCP server, APScheduler, BERTopic themes | Research-grade system |
12.4 Definition of Done per Phase
Each phase must pass all listed criteria before being marked complete:
P0 (Foundation):
schema.sqlcreates all tables without errors- Empty pipeline runs cleanly with exit code 0
oracle-pipeline.shis idempotent (safe to run twice)- Adapter registry loads all 6 adapters
P1 (First data):
- arXiv + RSS adapters fetch successfully
- Entries stored in SQLite with correct schema
- Keyword convergence detects at least 1 theme
- Deduplication works (no duplicate entries)
P2 (Full ingest):
- All 6 adapters fetch successfully in one run
- Rate limiting enforced per adapter
- Structured JSON logs emitted per pipeline stage
- No adapter failure kills the pipeline
P3 (Falsification):
- Exponential decay scoring implemented
- Ollama summarization works with graceful degradation
- Trend verdicts computed: confirmed/emerging/dead
- Running over 7 days shows false signals dying
P4 (Consumption):
- REST API endpoints return valid JSON
- Daily file drops written to
/output/ - Health endpoint reports accurate status
- Alerting webhooks fire on simulated failures
P5 (Validation):
- 7-day UAT: pipeline runs unattended without intervention
- Hermes cron integration works
- Exit codes correct: 0=success, 1=partial, 2=failure
- Pipeline completes <30 minutes end-to-end
P6 (Scale):
- sqlite-vec + embeddings operational
- MCP server responds to all 3 tools
- APScheduler handles per-source intervals
- System stays within 150MB pipeline memory budget
General: No critical bugs open, all unit tests pass, CI green, security scan clean.
13. Backport to PRD: Requirements to Add
The following requirements are implied by this design and should be added to docs/MVP-PRD.md:
13.1 Platform requirements (REQ-PLT-XX)
| ID | Requirement |
|---|---|
| REQ-PLT-05 | All processes run as standard user (no sudo) — already exists |
| REQ-PLT-10 | The pipeline process shall not exceed 500MB of RSS memory (excluding host-level Ollama) |
| REQ-PLT-15 | The system shall support deployment on a VPS with 2GB total RAM (pipeline + Ollama + OS) |
| REQ-PLT-20 | Ollama inference shall run as a host-level service, not inside the pipeline container |
| REQ-PLT-25 | The pipeline shall use SQLite as the sole database (no PostgreSQL, no Elasticsearch, no Redis) |
13.2 Reliability requirements (REQ-REL-XX)
| ID | Requirement |
|---|---|
| REQ-REL-05 | The system operates autonomously without human interaction — already exists |
| REQ-REL-10 | Failed source fetches retry with exponential backoff — already exists |
| REQ-REL-15 | Previously stored data is not lost on restart or failure — already exists |
| REQ-REL-20 | The pipeline completes successfully even if 1+ sources are unavailable — already exists |
| REQ-REL-25 | If Ollama is unreachable, ingestion continues and summarization defers to the next run |
| REQ-REL-30 | The pipeline uses flock/PID file to prevent overlapping runs |
13.3 Observability requirements (REQ-DIAG-XX)
| ID | Requirement |
|---|---|
| REQ-DIAG-05 | Structured JSON logs with timestamps and severity — already exists |
| REQ-DIAG-10 | Health endpoint reports system status and last successful run — already exists |
| REQ-DIAG-15 | Per-source success/failure and fetch counts logged per run — already exists |
| REQ-DIAG-20 | Webhook alert (Discord/Slack) fires when an adapter fails for 2+ consecutive days |
| REQ-DIAG-25 | run_log table captures per-run metrics queryable via SQL |
13.4 Integration requirements (REQ-INT-XX)
| ID | Requirement |
|---|---|
| REQ-INT-05 | Structured, machine-readable output (JSON) consumable by external tools — already exists |
| REQ-INT-10 | Adapter layer supports adding new sources without modifying core pipeline logic — already exists |
| REQ-INT-15 | REST API exposes GET /trends, /entries, /verdicts endpoints |
| REQ-INT-20 | Daily file drop at configurable path with structured output (YAML + JSON) |
| REQ-INT-25 | MCP server exposes oracle_search, oracle_trends, oracle_verdicts tools (Phase 6) |
13.5 Data requirements (REQ-DATA-XX) (new section)
| ID | Requirement |
|---|---|
| REQ-DATA-05 | Entries are deduplicated by (source, source_id) with cross-source URL hash deduplication |
| REQ-DATA-10 | FTS5 full-text index on title and extracted_text for keyword search |
| REQ-DATA-15 | Convergence detection: theme appears in ≥3 independent sources within 24h window |
| REQ-DATA-20 | Falsification uses exponential decay scoring (configurable λ), not fixed day thresholds |
| REQ-DATA-25 | Data retention: raw entries 90 days, summaries/convergence 365 days, periodic VACUUM |
13.6 Security requirements (REQ-SEC-XX)
| ID | Requirement |
|---|---|
| REQ-SEC-05 | All outbound HTTP uses TLS — already exists |
| REQ-SEC-10 | No secrets hardcoded or stored in plaintext — already exists |
| REQ-SEC-15 | Least-privilege access for outbound API calls — already exists |
| REQ-SEC-20 | SQLite database file permissions set to 600 (owner-only read/write) |
| REQ-SEC-25 | All SQL queries use parameterized statements (no string concatenation) |
13.7 Scheduling requirements (REQ-SCH-XX) (new section)
| ID | Requirement |
|---|---|
| REQ-SCH-05 | Pipeline entry point is oracle-pipeline.sh (idempotent, single command) |
| REQ-SCH-10 | Default schedule: 13:00 UTC daily |
| REQ-SCH-15 | Exit codes: 0 = success, 1 = partial failure, 2 = total failure |
| REQ-SCH-20 | Overlapping run prevention via flock or PID file check |
14. Design Decisions Summary (Why)
| Decision | Why | Revisit |
|---|---|---|
| SQLite over PostgreSQL | Zero external dependency, single file, FTS5 built-in, handles 1M rows fine. pgvector is premature at this scale. | P6 (if scale demands pgvector) |
| Ollama host-level | 150MB container cannot fit Ollama + model (~2GB). Host-level lets pipeline stay within budget. | P2 (if inference needs change) |
| Flask over FastAPI | ~1MB vs ~100MB runtime overhead. FastAPI is Phase 2 target; Flask suffices for internal REST API. | P2 (when FastAPI becomes viable) |
| Keyword co-occurrence (Phase 1) | Zero-dependency, explainable, works at 150MB. Embeddings (Phase 2) add semantic convergence. | P2 (when embeddings ready) |
| Fixed themes + catch-all | BERTopic requires 4GB RAM. Fixed themes with "other" bucket is the pragmatic constraint choice. | P6 (when auto-discovery needed) |
| Cron over APScheduler (Phase 1) | OS-level, zero process memory cost. APScheduler is Phase 2 for dynamic scheduling. | P2 (if per-source intervals needed) |
| HTTP-only adapters | All 6 sources have programmatic APIs. Playwright adds 300MB+ overhead and fragility. | N/A (stable) |
| Exponential decay over 7-day rule | One-line formula, no fixed threshold. Handles fast-dying and slow-burn trends naturally. | N/A (stable) |
| Deduplication required | arXiv papers appear on HN/Reddit/Twitter. Without dedup, same signal counted 3× = false convergence. | N/A (stable) |
| Graceful degradation on Ollama | Ingestion must not depend on summarization. Store raw data, defer summaries. | N/A (stable) |
15. Testing Strategy
15.1 Unit Tests
- Per adapter: Test each of the 6 adapters with known-good endpoints. Verify:
- Returns valid JSON with required schema fields (source, source_id, title, url, timestamp)
- Handles rate limits gracefully (no infinite loops)
- Correct error codes for 429/503
- Scoring functions: Unit tests for exponential decay, convergence scoring, and theme matching. Include edge cases (empty input, negative scores).
15.2 Integration Tests
- Run pipeline against seed dataset (100 entries from arXiv + RSS). Verify:
- All adapters fetch successfully
- Deduplication removes duplicates correctly
- Theme detection identifies at least 3 themes
- No critical errors in logs
15.3 E2E Tests
- Cron-to-snapshot: Run full pipeline via cron, verify output files match expected snapshot (golden file comparison)
- Stress test: Run 10 consecutive daily cycles with simulated failures (adapter timeout, Ollama down). Verify graceful degradation and recovery
Test coverage goal: 80% of critical paths covered by automated tests. Manual testing for theme quality and summary accuracy.
16. Deployment and CI/CD
16.1 CI Pipeline
Updated 2026-07-12: CI originally built a Docker image per push. The live system has no CI/Docker build. The equivalent checks run manually/optionally: lint,
pip audit, andpytestif tests exist. Container build is deferred.
On every push to MVP-milestone:
- Lint:
ruff check,mdlint docs/ - Test: Run unit tests (
pytest tests/), integration tests with seed dataset - Scan: Run
pip audit; fail if critical vulnerabilities (>CVSS 7.0) found - (Deferred: Docker image build/tag with commit SHA)
16.2 Deployment
Deployment reality (updated 2026-07-12): This document originally specified Docker containerization (150MB cap, non-root,
python:3.11-slim). The deployed system does not use containers. It runs as bare Python scripts under thevpsadminstandard user, invoked by Hermes cron (oracle-pipeline.shat 13:00 UTC). The 150MB budget, non-root, and read-only-filesystem hardening described below remain design intent and are tracked as a future containerization task — not the current state. Ollama is also not deployed; summarization is deferred (graceful degradation is the live behavior).
- Local dev / Production: Bare Python on the VPS, no Docker.
pip installdeps, runschema.sql, source.env, then Hermes cron drivesoracle-pipeline.sh. - Ollama: Not currently running. Pipeline stores raw text and defers summarization (see §10.2 / REQ-REL-25). Adding Ollama is a separate step, host-level.
16.3 Rollback Procedure
If deployment fails or quality degrades:
- Stop services:
systemctl stop oracle-pipeline qwythos-gpu0 - Restore previous code:
git checkout <good-tag> - Rebuild Docker image from restored code
- Restart services:
systemctl start oracle-pipeline qwythos-gpu0 - Verify health:
curl http://localhost:8081/v1/health
Rollback window: Must complete within 5 minutes of failure detection.
Tagging: Each deploy is tagged with semantic versioning (v1.0.0, v1.1.0) for easy rollback reference.
17. Operational Runbooks
17.1 Daily Pipeline Verification
At 13:00 UTC after each run:
- Check logs:
journalctl -u oracle-pipeline --since "today" | grep ERROR - Verify exit code in
/var/log/oracle-pipeline/run_log.txt— should be 0 - Confirm output:
ls -lh /output/$(date +%Y-%m-%d)/— should havesummary.jsonandmetrics.json - If any check fails, investigate with the relevant runbook below
17.2 Database Recovery from Corruption
If sqlite3 oracle.db 'PRAGMA integrity_check' returns errors:
- Stop pipeline:
systemctl stop oracle-pipeline - Restore from last backup:
cp /backup/oracle.db.bak-YYYYMMDD oracle.db - Verify integrity:
sqlite3 oracle.db 'PRAGMA integrity_check'— should return "ok" - Restart pipeline:
systemctl start oracle-pipeline
Prevention: Daily compressed backups to /backup/, retention 30 days.
17.3 Adapter Failure Investigation
If adapter fails repeatedly (>2 consecutive runs):
- Check logs:
journalctl -u oracle-pipeline --since "today" | grep -A5 "Adapter" - Test endpoint manually:
curl -X GET <adapter_url>— verify HTTP status - Check rate limit headers:
curl -I <adapter_url> | grep -i 'x-ratelimit' - If rate-limited: Wait
retry_afterseconds, pipeline retries next cycle - Escalate to project lead if >3 consecutive failures
17.4 Ollama Service Restart
If Ollama becomes unresponsive:
- Check status:
systemctl status qwythos-gpu0.service - View logs:
journalctl -u qwythos-gpu0.service --since "today" - Restart service:
systemctl restart qwythos-gpu0.service - Verify health:
curl http://localhost:8081/v1/health— should return{ "status": "ready" } - Check GPU:
nvidia-smi— ensure service actually loaded
Escalation: If issue persists after restart, notify project lead.
Runbook maintenance: Update runbooks when features change. Document changes in runbook_changes.md.
18. Risk Register and Assumptions
18.1 Risk Register
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Source API changes (rate limits, endpoint shifts) | High | High | Auto-retry with exponential backoff; monitor adapter health; log API changes for review |
| Model quality drift (summarization accuracy degrades) | Medium | Medium | Weekly golden sample comparison; rollback if hallucination rate >10% or summary <50 tokens |
| Disk space exhaustion (backups, logs) | High | High | Automated 30-day retention; alert at 85% usage; compress old backups |
| Single operator bottleneck (manual theme review) | Medium | Medium | Documented theme governance; quarterly review cycle; catch-all "other" bucket |
| Network outage (all adapters fail) | Low | High | Graceful degradation: store raw data without summaries; resume when network returns |
| SQLite corruption during schema migration | Low | Critical | Pre-migration backup; atomic migration scripts; integrity check after each migration |
18.2 Key Assumptions
- All 6 source APIs remain stable for at least 6 months (no breaking changes)
- GPU memory available for Qwythos inference (~15GB free on GPU0)
- Network connectivity to all sources is available during pipeline runs
- Single operator can complete manual theme review within 24 hours
- Disk space sufficient for 30-day backup retention (~10GB)
Risk review: Quarterly review of this register; update mitigations when new risks emerge.
Assumption tracking: If an assumption proves false, document the deviation in assumption_deviations.md and reassess risks.
19. Dependencies and Tooling
19.1 Python Runtime Dependencies
All dependencies pinned to exact versions in requirements.txt:
- requests: HTTP client for all source APIs
- beautifulsoup4: HTML parsing for RSS feeds
- feedparser: RSS/Atom feed handling (primary)
- tenacity: Retry logic with exponential backoff (used by all adapters)
- flask: Internal REST API for Phase 4+ (served on port 8081)
No external database dependencies — SQLite is built into Python.
19.2 Runtime System Dependencies
- Ollama: Host-level inference service (~2GB RAM, GPU acceleration)
- systemd: Service management for pipeline and inference (
oracle-pipeline.service,qwythos-gpu0.service) - cron: Daily schedule trigger (Phase 1); systemd timers for Phase 2
- sqlite3: Database CLI for backup/recovery commands
19.3 Development Tooling
- ruff: Linting and formatting (Python)
- pytest: Unit and integration test framework
- docker: Containerization for pipeline
- pip audit / safety: Dependency vulnerability scanning (CI checks)
19.4 Dependency Update Policy
- Critical security updates: Apply within 7 days of CVE disclosure
- Minor version updates: Test in staging before production deployment
- Major version upgrades: Require migration scripts and rollback plan
Tooling policy: No new dependencies without approval from project lead; document rationale in dependency_justifications.md.
Lock file: requirements.txt pinned to exact versions for reproducible builds.
Document prepared via multi-model analysis: Qwythos-9B (architectural critique), Qwen3.5-9B (implementation evaluation), and cross-review synthesis. 4 delegations, 2 rounds of debate. Raw reviews saved in the same directory.