Merge MVP-milestone docs into main #15

Merged
Ty merged 14 commits from MVP-milestone into main 2026-07-12 04:12:12 +00:00
6 changed files with 1238 additions and 0 deletions
+42
View File
@@ -0,0 +1,42 @@
# Deployment Plan — Athena MVP
## Prerequisites
- Docker installed on the target host
- Network access to the 6 source APIs/feeds (arxiv, github, huggingface, hackernews, reddit, rss_feeds)
- `GITHUB_TOKEN` available as an environment variable (optional, but raises the GitHub rate limit from 60/hr to 5000/hr)
- `HUGGINGFACE_TOKEN` available as an environment variable
- A local Ollama instance running `llama3.2:1b`, or an alternate reachable inference backend if swapping — per the model-agnostic `summarize(text) -> (summary, model)` contract
- Hermes cron infrastructure available and able to invoke `oracle-pipeline.sh`
## Setup
1. Clone `main` (not a milestone branch) onto the target host
2. `pip install -r requirements.txt` if present; otherwise confirm stdlib + `requests` are available
3. Run `schema.sql` against a fresh `oracle.db` — this file is git-ignored and created locally, never committed
4. Export required environment variables (`GITHUB_TOKEN`, `HUGGINGFACE_TOKEN`) — never hardcode these
5. Build and run inside Docker with the 150MB memory cap and non-root user enforced, per the PRD platform requirements
6. Manual smoke test: run `python3 pipeline.py` once and confirm ingest → store → summarize → score completes without errors before handing off to cron
## Cron / Scheduling
1. Confirm `oracle-pipeline.sh` is the entry point Hermes cron calls
2. Schedule for 13:00 UTC daily
3. Add a lock file or PID check so overlapping runs can't happen if a prior run is still in progress
4. Confirm the cron environment actually carries the exported tokens — cron environments are frequently minimal and won't inherit an interactive shell's exports
## Hermes Integration
1. Confirm the wrapper script's exit codes are meaningful (0 = success, non-zero = failure) so Hermes can act on them
2. Define where Hermes should look for pipeline output/logs
3. Decide on a failure-notification path (Hermes alert, log flag, etc.) — **not yet specified, needs a decision**
## Monitoring
1. Aggregate logs from each pipeline stage (ingest, store, summarize, score)
2. Track theme-scan new-arrival counts per cycle per theme — this is the core signal the falsification logic depends on, so it deserves visibility beyond raw logs
3. Add a heartbeat/dead-man's-switch alert if a scheduled run doesn't fire, rather than relying on someone noticing missing data days later
4. Watch memory usage against the 150MB cap under real production load, not just dev conditions
## Rollback
1. Back up `oracle.db` before any schema change — it's git-ignored and not recoverable from the repo itself
2. If a bad deploy breaks the pipeline, revert to the last known-good commit on `main` and redeploy the Docker image
3. Check falsification state (new-arrivals counters) after any rollback — rolling back mid-window could distort the 7-day dead-thesis calculation if not handled carefully
---
*Draft prepared by Claude from the README, whitepaper falsification logic, and MVP-PRD platform requirements on `main`. The Hermes failure-notification path is the one open decision blocking this from being final.*
+849
View File
@@ -0,0 +1,849 @@
# Athena-Oracle: Development Design Document
**Version:** 0.1.0
**Date:** 2026-07-08
**Status:** Draft — pre-implementation
**Source branch:** `MVP-milestone`
---
## 1. North Star
> *Surfacing cross-source convergence and using falsification to distinguish real momentum from noise.*
Athena is an autonomous research intelligence engine that ingests from multiple fragmented sources, detects when the same signals appear across independent channels, and uses decay-based falsification to separate genuine trends from one-day spikes. The system is model-agnostic, lightweight, and designed to run unattended on a resource-constrained VPS.
### How this design supports the north star
| North Star Principle | Design Decision | Why |
|---|---|---|
| Cross-source convergence | Keyword co-occurrence matrix (Phase 1), embeddings + vector search (Phase 2) | Keyword co-occurrence is the simplest convergence detector: if the same entity appears in ≥3 independent sources within a time window, it's converging. Embeddings Phase 2 adds semantic convergence for signals that use different words but mean the same thing. |
| Falsification over confirmation | Exponential decay scoring | A 7-day hard cutoff is blunt: some trends die in 48 hours, some take 60 days to validate. `score = base_score * e^(-λ * days_since_last_signal)` naturally scores dying trends low and sustained trends high without arbitrary day thresholds. |
| Autonomous operation | Cron/scheduled pipeline + graceful degradation | The pipeline runs daily without human intervention. If Ollama is down, ingestion continues and summarization defers to the next run. If one adapter fails, the rest still run. |
| Lightweight deployment | Python + SQLite + Flask + host-level Ollama | No PostgreSQL, no Elasticsearch, no Redis, no Kubernetes. A single Python process, a single SQLite file, and an external Ollama REST API. |
### MVP Success Criteria
The MVP is considered successful when:
1. **Research Time**: Bob answers 'what's trending this week' in <15 min of research time total (across all sources).
2. **Signal Detection**: Over 7 days, the theme scan correctly identifies ≥1 real trend that has genuine cross-source convergence.
3. **Noise Rejection**: Over the same 7-day period, the falsification engine kills ≥1 false signal (a one-day spike with no sustained arrivals) to demonstrate decay-based filtering is working.
---
## 2. System Architecture
```
┌─────────────────────────────────────────────────────┐
│ Athena-Oracle Pipeline (Python, ~150-500MB) │
│ │
│ ┌────────────┐ ┌─────────┐ ┌──────────────────┐ │
│ │ GitHub │ │ arXiv │ │ Reddit │ │
│ │ Adapter │ │ Adapter │ │ Adapter │ │
│ └──────┬─────┘ └────┬────┘ └────────┬─────────┘ │
│ │ │ │ │
│ ┌──────┴─────────────┴──────────────┴──────────┐ │
│ │ Deduplication (URL hash + title simhash) │ │
│ └────────────────────┬─────────────────────────┘ │
│ │ │
│ ┌────────────────────┴─────────────────────────┐ │
│ │ Theme Tagging │ │
│ │ Phase 1: Keyword co-occurrence + fixed seeds│ │
│ │ Phase 2: all-MiniLM embeddings + BERTopic │ │
│ └────────────────────┬─────────────────────────┘ │
│ │ │
│ ┌────────────────────┴─────────────────────────┐ │
│ │ Falsification Engine │ │
│ │ Exponential decay scoring per theme │ │
│ │ Convergence threshold: ≥3 independent sources│ │
│ └────────────────────┬─────────────────────────┘ │
│ │ │
│ ┌────────────────────┴─────────────────────────┐ │
│ │ SQLite │ │
│ │ entries table + run_log + FTS5 index │ │
│ │ Phase 2: + sqlite-vec extension │ │
│ └──────────────────────────────────────────────┘ │
│ │
│ Output layer: │
│ ┌──────────┐ ┌──────────────┐ ┌────────────────┐ │
│ │ Flask API│ │ File drops │ │ MCP server │ │
│ │ (Phase 4)│ │ (Phase 4) │ │ (Phase 6) │ │
│ └──────────┘ └──────────────┘ └────────────────┘ │
│ │
│ Structured JSON logging → stdout + file rotation │
│ Discord/Slack webhook on 2+ day adapter failure │
└─────────────────────────────────────────────────────┘
┌───────────────────────────────┘
│ HTTP REST API
┌─────────────────────┐
│ Ollama (host-level) │
│ llama3.2:1b │
│ (summarization) │
└─────────────────────┘
```
### Memory budget
| Component | Phase 1 | Phase 2 |
|---|---|---|
| Python runtime + deps | ~80 MB | ~80 MB |
| SQLite (in-process) | ~10 MB | ~10 MB |
| all-MiniLM embeddings | — | ~80 MB |
| sqlite-vec extension | — | ~2 MB |
| Flask | ~1 MB | ~1 MB |
| **Pipeline total** | **~90 MB** | **~173 MB** |
| Ollama + model (host-level) | ~2 GB | ~2 GB |
| **System total** | **~2.1 GB** | **~2.2 GB** |
The 150MB constraint applies to the pipeline process. The full system footprint including Ollama is ~2GB.
### Daily Run Flow
1. **Preflight** — Check disk space, verify DB integrity (`PRAGMA integrity_check`), load adapter config
2. **Ingest** — Run all 6 adapters in parallel, collect raw items per source
3. **Dedup** — Hash URLs, match against existing entries, insert only new items
4. **Theme Tag** — Run keyword co-occurrence against new items, tag themes
5. **Falsification** — Recompute decay scores for all active themes, kill dead theses
6. **Archive** — Move entries older than 90 days to archive table
7. **Report** — Write daily summary to `/output/`, update run_log
### Failure Modes
| Component | Failure | Impact | Recovery |
|---|---|---|---|
| Adapter | Rate limit / 503 | Missing items from that source | Retry next cycle; pipeline continues |
| Ollama | Down | Summaries skipped | Entries stored with `summary = null`, deferred to next run |
| SQLite | Disk full | No writes | Alert via webhook; manual cleanup |
| SQLite | Corruption | Data loss | Restore from last backup |
| Network | Outbound blocked | All adapters fail | Alert; pipeline exits code 2 |
---
## 3. Data Layer
### 3.1 SQLite schema
Core tables (from `schema.sql`):
- **entries** — one row per ingested item, deduplicated by `(source, source_id)`
- **run_log** — one row per pipeline run, tracks per-source success/failure
- **FTS5 virtual table** — full-text search over `title` and `extracted_text`
- **Phase 2: sqlite-vec** — vector index for semantic similarity queries
### 3.2 Why SQLite
- Zero external dependency, single file, survives container restarts
- FTS5 is built-in (no separate search engine)
- Handles 100K-1M rows without performance issues
- sqlite-vec extension adds vector search without a separate database
- No connection pooling needed (single-writer pipeline)
- Postgres/pgvector is premature optimization at this scale
### 3.3 Data retention
- Raw entries: 90 days
- Summaries and convergence scores: 365 days
- Periodic `VACUUM` to reclaim space
- `archive.py` handles cold storage rotation (deferred to Phase 3)
### 3.4 Backup Strategy
- **Pre-schema-change backup**: Before any `ALTER TABLE` or schema modification:
```bash
sqlite3 oracle.db '.backup oracle.db.bak'
```
Store `.bak` files with date suffix in `/backup/` (`oracle.db.bak-YYYYMMDD`).
- **Daily compressed backup**: At 01:00 UTC (off-peak):
```bash
tar -czf /backup/oracle.db.$(date +%Y%m%d).tar.gz oracle.db
```
Retain 30 days of backups; purge older: `find /backup -name '*.tar.gz' -mtime +30 -delete`
- **Recovery**: Restore from backup with `cp /backup/oracle.db.bak-YYYYMMDD oracle.db`, verify with `PRAGMA integrity_check`
### 3.5 Schema Migration
- Versioned migration files in `migrations/` directory (e.g., `001_initial_schema.sql`, `002_add_theme_tags.sql`)
- Applied at startup: pipeline checks `schema_version` table, runs any unapplied migrations in order
- Each migration is a single atomic SQL file; no partial migrations
- Rollback: each migration includes a comment with the reverse SQL if needed
---
## 4. Adapter Layer
### 4.1 Source adapters (HTTP-only)
| Adapter | API | Rate limit | Auth required |
|---|---|---|---|
| GitHub | REST API | 60/hr (unauth), 5000/hr (token) | `GITHUB_TOKEN` |
| arXiv | REST API | 1 req/sec (polite) | No |
| Reddit | RSS/JSON | ~100 req/min | No (but OAuth recommended) |
| Hacker News | Firebase API | Unofficial, ~30 req/sec | No |
| HuggingFace | REST API | Throttled if aggressive | `HUGGINGFACE_TOKEN` |
| RSS Feeds | RSS XML | Varies | No |
**Decision:** HTTP-only adapters, no Playwright/Selenium. All 6 sources have programmatic APIs. Playwright would add Chromium's 300MB+ overhead and fragility.
### 4.2 Deduplication
arXiv papers appear on HN, Reddit, and Twitter. Without deduplication, the same signal is counted 3× and produces false convergence.
- **Phase 1:** UNIQUE constraint on `(source, source_id)` + URL hash dedup across sources
- **Phase 2:** SimHash/MinHash content fingerprinting for near-duplicate detection
### 4.3 Rate limiting and retries
- Per-adapter rate limits enforced in the adapter class
- `tenacity` library for exponential backoff on transient failures (429, 503, timeout)
- One failing adapter does not kill the pipeline
### 4.4 Adapter Interface Contract
All adapters must implement this minimal interface:
- **name() → str**: Unique identifier for the source (e.g., `"arxiv"`)
- **fetch(query, limit) → list[dict]**: Returns normalized items with required schema fields:
- `source` (str): Source name matching `name()`
- `source_id` (str): Unique per-source ID for deduplication
- `title` (str): Human-readable title
- `url` (str): Direct URL to the entry
- `timestamp` (datetime): Publication/update time
- `raw_score` (float): Source-specific signal strength (0.010.0)
- `body` (str): Raw text content for summarization and theme tagging
**Custom Exceptions:**
- `RateLimitError`: Raised when source returns 429 or similar; includes `retry_after` (seconds)
- `SourceUnavailableError`: Raised when source is down (5xx) or unreachable
Adapters must not raise other exceptions on normal operation; unexpected errors should be logged with full traceback and surfaced via the run_log, not propagated to the pipeline.
---
## 5. Theme Tagging and Convergence Detection
### 5.1 Phase 1: Keyword co-occurrence
Pre-defined keyword dictionaries per theme. An entry is tagged if ≥2 keywords from a theme dictionary appear in its title or extracted text. A theme "converges" if it appears in ≥3 independent sources within the last 24 hours.
**Why keyword first at 150MB:** Keyword matching is zero-dependency, explainable, and works within the memory constraint. FTS5 provides fast retrieval.
**Fixed themes:** The initial 4 themes (`tool-call`, `context`, `compute`, `trust`) are seeds, not a hard limit. An "other" catch-all bucket captures signals that don't match predefined themes.
### 5.2 Phase 2: Embeddings + auto-discovery
- **all-MiniLM-L6-v2** (22M params, ~80MB) for sentence embeddings
- **sqlite-vec** for in-database ANN search
- **BERTopic** (or equivalent) for semi-supervised theme discovery, seeded from the Phase 1 dictionary
- Hybrid query: FTS5 for precision (keyword match) + vector for recall (semantic match), merged via Reciprocal Rank Fusion
**Why not keyword forever:** Keyword matching cannot detect semantic convergence (different words, same concept) and requires constant manual dictionary updates. Embeddings are the eventual target; Phase 1 is the bridge.
### 5.3 Convergence scoring
```
convergence_score = Σ(source_weights) × temporal_proximity × theme_entropy
where:
source_weights: arXiv=2.0, GitHub=1.5, HN=1.0, Reddit=0.8, HF=1.2, RSS=0.5
temporal_proximity: e^(-0.1 * hours_since_first_signal)
theme_entropy: log2(number_of_independent_sources)
```
Thresholds:
- `≥ 3.0` → "confirmed" trend
- `≥ 1.5` → "emerging" signal
- `< 1.5` → "noise"
### 5.4 Theme Governance
Themes are defined in a single YAML file: `themes.yaml`. Each theme entry includes:
- `name`: human-readable identifier (e.g., `tool-call`)
- `keywords`: list of keyword patterns for Phase 1 regex matching
- `owner`: responsible person (for quarterly review)
- `created_at`: ISO date of creation
- `status`: `active` or `deprecated`
**Lifecycle:**
- **Propose**: New themes require owner nomination + approval from project lead
- **Review**: Active themes are reviewed quarterly; deprecated if <2 hits in 30 days
- **Retire**: Deprecated themes are excluded from convergence scoring after 90 days
- **Reinstate**: Deprecated themes can be reactivated if signals re-emerge
The "other" catch-all bucket captures signals that don't match any active theme and is reviewed during quarterly theme audits for potential new theme creation.
---
## 6. Falsification Engine
### 6.1 Exponential decay scoring
Replace the 7-day dead thesis rule with:
```
thesis_score = initial_score × e^(-λ × days_since_last_signal)
where λ = 0.1 (configurable)
```
A thesis is "dead" when its score falls below a configurable threshold (default: 0.1), not when it hits a fixed day count. This naturally handles:
- Fast-dying trends (score drops quickly)
- Slow-burn trends (score stays elevated)
- Revived trends (new signal resets the decay clock)
### 6.2 Cross-source validation
A signal is flagged "unverified" if:
- Only 1 source has primary (non-derivative) coverage
- The signal appears only in echo chambers (e.g., HN upvotes ≠ real adoption)
- A counter-narrative exists in the same time window
### 6.3 Calibration Process
- **Initial parameters**: Start with λ = 0.1 (half-life ~7 days) for all themes at deployment
- **Validation window**: After 30 days of live operation, validate against historical data
- **Adjustment triggers**:
- If >20% of confirmed real trends were falsely killed → decrease λ (e.g., to 0.05, slower decay)
- If >30% of noise signals were incorrectly confirmed → increase λ (e.g., to 0.15, faster decay)
- **Documentation**: Record calibration decisions in `calibration_log.md` with date, old/new λ values, and rationale
Re-calibrate quarterly or whenever a major theme dictionary change is made.
---
## 7. Output Layer
### 7.1 Consumer interfaces (progressive rollout)
| Consumer | Interface | Phase |
|---|---|---|
| **Bob** (trend tracker) | Flask REST API: `GET /trends?theme=&period=7d` | 4 |
| **Alice** (content creator) | Daily file drops: `/output/YYYY-MM-DD/trends.yaml` | 4 |
| **Sam** (Hermes agent) | MCP server: `oracle_search`, `oracle_trends`, `oracle_verdicts` | 6 |
### 7.2 REST API (Phase 4)
Flask endpoints:
- `GET /health` — pipeline status, last run time, adapter health
- `GET /trends` — active themes with convergence scores
- `GET /entries` — search entries (keyword + phase 2: semantic)
- `GET /verdicts` — confirmed/dead theses
- `GET /convergence` — cross-source convergence matrix
### 7.3 File drops (Phase 4)
Daily structured output at a known path:
```
/output/YYYY-MM-DD/
trends.yaml # Human-readable daily digest
signals.json # Structured machine-readable output
verdicts.json # Confirmed/dead thesis list
```
### 7.4 MCP server (Phase 6)
MCP tools for Hermes agent integration:
- `oracle_search(query, source, date_range)` — search entries
- `oracle_trends(theme, convergence_threshold)` — get active trends
- `oracle_verdicts(status)` — confirmed or dead theses
- `oracle_latest(source)` — most recent entry per source
### 7.5 MCP Tool Signatures
All MCP tools must implement these minimum signatures:
- **get_trends()**:
- Request: `{}` (no params)
- Response: `{ "trends": [{"name": str, "score": float, "sources": [str], "decay_score": float}] }`
- **search_entry(query: str, source: str | None = None)**:
- Request: `{ "query": str, "source": str | null }`
- Response: `{ "entries": [{"title": str, "url": str, "summary": str, "score": float}] }`
- **get_convergence_report()**:
- Request: `{}`
- Response: `{ "converged": [{"entity": str, "sources": [str], "confidence": float}] }`
Tools must validate input types and return empty arrays (not errors) for valid queries that yield no results.
### 7.6 API Authentication Model
- **Phase 4 (MVP)**: Simple API key in `X-API-Key` header. No expiration, stored in config file (`/etc/athena/api_keys.yaml`)
- **Phase 6 (Production)**: JWT bearer token with scopes (read-only, read-write, admin). Tokens expire after 24 hours; refresh via `/auth/token` endpoint
Auth failures return HTTP 401 with `{ "error": "unauthorized" }`. Rate limiting applies per-key: 100 requests/minute.
---
## 8. Observability and Reliability
### 8.1 Logging
Structured JSON logging via stdlib `logging` with JSON formatter. Per-pipeline-stage logs (ingest, dedup, theme, falsification, summarize) with source-level granularity.
### 8.2 Health endpoint
`GET /health` returns:
```json
{
"status": "ok",
"last_run": "2026-07-08T13:00:00Z",
"last_run_duration_sec": 245,
"entries_since_last_run": 127,
"adapters": {
"github": {"status": "ok", "fetched": 20},
"arxiv": {"status": "ok", "fetched": 15},
"reddit": {"status": "error", "fetched": 0, "error": "429 rate limited"}
}
}
```
### 8.3 Alerting
Discord/Slack webhook triggered when:
- An adapter fails for 2+ consecutive days
- Pipeline run exceeds 2× expected duration
- SQLite database integrity check fails
### 8.4 Graceful degradation
If Ollama is unreachable:
- Ingestion continues normally
- Summarization is skipped, entries stored with `summary = null`
- **Deferral**: next run summarizes pending entries
- **Alert**: "summarization deferred, N entries pending"
### 8.5 Log Retention and SLIs
- **Log retention**: 30 days rolling; gzip-compressed after 7 days to save disk space
- **SLI definitions**:
- **Pipeline success rate**: >95% of daily runs complete without critical failure (exit code 2)
- **Adapter availability**: >90% of scheduled runs successfully fetch each source (per-source metric)
- **Theme detection accuracy**: ≥80% of manually verified trends identified correctly in first week
### 8.6 Alerting Matrix
| Condition | Channel | Severity | Response Time |
|---|---|---|---|
| Adapter fails >2 consecutive runs | Discord webhook | P2 | Investigate within 1 hour |
| Pipeline exit code 2 | Discord webhook + email | P1 | Investigate within 30 minutes |
| DB disk usage >85% | Discord webhook | P2 | Investigate within 2 hours |
| Ollama unreachable >5 min | Discord webhook | P2 | Restart service if needed |
| Pipeline success rate <90% for 3 days | Email + dashboard | P3 | Review next cycle |
Alerts are deduplicated: same condition won't fire again until resolved.
---
## 9. Scheduling
### 9.1 Phase 1: Cron
- `oracle-pipeline.sh` invoked by cron at 13:00 UTC daily
- `flock`/PID file prevents overlapping runs
- Exit codes: 0 = success, 1 = partial failure, 2 = total failure
### 9.2 Phase 2: systemd timers
- `Persistent=true` catches up on missed runs
- `RandomizedDelaySec` prevents thundering herd
- `OnFailureSec` for retry logic
- Better logging than cron (`journalctl -u athena-timer`)
### 9.3 Why not APScheduler (Phase 1)
APScheduler adds in-process async daemon overhead. Cron/systemd is OS-level, zero process memory cost, and sufficient for daily runs. APScheduler is the Phase 2 target if dynamic scheduling (user-configurable refresh rates, per-source intervals) is needed.
### 9.4 Scheduling Recommendation
- **Phase 1 (MVP)**: Use system cron for daily runs at 13:00 UTC. Sufficient for fixed schedule, zero process memory cost, OS-level reliability.
- **Phase 2**: Switch to systemd timers if dynamic scheduling needed (skip runs on holidays, adjust time zones). Better observability and integration with monitoring tools.
- **APScheduler**: Only use if per-source intervals are required (e.g., arXiv every 6 hours, Reddit every 30 min). Adds ~50MB process overhead — not recommended for Phase 1.
**Recommendation**: Start with cron for Phase 1. Evaluate whether Phase 2 sources need different intervals before considering systemd timers or APScheduler.
---
## 10. Inference
### 10.1 Summarization
**Model:** Ollama `llama3.2:1b` (or `qwen2.5:0.5b` for lower resource)
**Deployment:** Host-level Ollama service, pipeline calls via HTTP REST API
**Contract:** Model-agnostic — `summarize(text) → (summary, model)` interface
**Graceful degradation:** If Ollama is down, store raw text and defer summarization
### 10.2 Why not in-container Ollama
Ollama daemon + 1B model requires ~2GB RAM. Running it inside the 150MB container is physically impossible. Running it host-level means the pipeline process stays within budget and Ollama can share resources with other services.
### 10.3 Model Upgrade Process
- **Swap model**: Replace model file; update systemd unit `ExecStart` path if needed
- **Restart service**: `systemctl restart qwythos-gpu0.service` (or equivalent unit)
- **Verify health**: `curl http://localhost:8081/v1/health` — should return `{ "status": "ready" }`
- **Quality check**: Run first summarization on known-good source (e.g., arXiv paper), compare output against baseline summary for content accuracy, length consistency, and hallucination rate
- **Rollback**: If quality degrades (e.g., summary length <50 tokens, hallucination rate >10%), revert to previous model file immediately
**Quality metrics**: Pass if summary is >50 tokens, no hallucinations on known entities, and theme detection matches golden samples. Only proceed with upgrade after successful verification.
---
## 11. Security
| Requirement | Implementation |
|---|---|
| No hardcoded secrets | `GITHUB_TOKEN`, `HUGGINGFACE_TOKEN` as env vars or mounted secret files |
| TLS for outbound | All HTTP adapters use `https://` |
| Least privilege | Pipeline runs as standard user (no sudo) |
| DB protection | `chmod 600 oracle.db` |
| Input sanitization | Parameterized SQL queries, no string concatenation |
### 11.2 MVP Security Baseline
- **Container hardening**:
- Run as non-root user (`user: nobody` in Dockerfile)
- Read-only filesystem where possible (except `/tmp`, `/var/log`)
- No SSH access inside container; pipeline is cron-triggered, no interactive access needed
- Minimal base image: `python:3.11-slim` (no dev tools, no git)
- **Dependency scanning**:
- CI pipeline runs `pip audit` or `safety check` on every push to `MVP-milestone`
- Fail build if critical vulnerabilities found (>CVSS 7.0)
- Warn on medium/high vulnerabilities; require manual review before merging
- **Secret management**:
- No secrets in code or config files (use environment variables at runtime)
- API keys stored in `/etc/athena/secrets.yaml` with restricted permissions (`chmod 0600`)
**Enforcement**: Security checks are automated in CI; local development is permissive but container builds must pass all scans.
---
## 12. Implementation Phases
| Phase | Scope | Deliverable |
|---|---|---|
| **P0: Foundation** | schema.sql + sqlite-vec design, adapter registry, oracle-pipeline.sh skeleton | Empty but valid pipeline |
| **P1: First data** | arXiv + RSS adapters, SQLite storage, keyword convergence, dedup | Live data flowing |
| **P2: Full ingest** | GitHub, HN, HF, Reddit adapters, rate limiting, structured logging | All 6 sources live |
| **P3: Falsification** | Exponential decay scoring, Ollama summarization, graceful degradation | Trend verdicts working |
| **P4: Consumption** | Flask API, daily file drops, health endpoint, alerting webhooks | Bob and Alice can consume |
| **P5: Validation** | 7-day UAT window, Hermes cron integration, exit codes | System runs unattended |
| **P6: Scale** | sqlite-vec + embeddings, MCP server, APScheduler, BERTopic themes | Research-grade system |
### 12.4 Definition of Done per Phase
Each phase must pass all listed criteria before being marked complete:
**P0 (Foundation)**:
- `schema.sql` creates all tables without errors
- Empty pipeline runs cleanly with exit code 0
- `oracle-pipeline.sh` is idempotent (safe to run twice)
- Adapter registry loads all 6 adapters
**P1 (First data)**:
- arXiv + RSS adapters fetch successfully
- Entries stored in SQLite with correct schema
- Keyword convergence detects at least 1 theme
- Deduplication works (no duplicate entries)
**P2 (Full ingest)**:
- All 6 adapters fetch successfully in one run
- Rate limiting enforced per adapter
- Structured JSON logs emitted per pipeline stage
- No adapter failure kills the pipeline
**P3 (Falsification)**:
- Exponential decay scoring implemented
- Ollama summarization works with graceful degradation
- Trend verdicts computed: confirmed/emerging/dead
- Running over 7 days shows false signals dying
**P4 (Consumption)**:
- REST API endpoints return valid JSON
- Daily file drops written to `/output/`
- Health endpoint reports accurate status
- Alerting webhooks fire on simulated failures
**P5 (Validation)**:
- 7-day UAT: pipeline runs unattended without intervention
- Hermes cron integration works
- Exit codes correct: 0=success, 1=partial, 2=failure
- Pipeline completes <30 minutes end-to-end
**P6 (Scale)**:
- sqlite-vec + embeddings operational
- MCP server responds to all 3 tools
- APScheduler handles per-source intervals
- System stays within 150MB pipeline memory budget
**General**: No critical bugs open, all unit tests pass, CI green, security scan clean.
---
## 13. Backport to PRD: Requirements to Add
The following requirements are implied by this design and should be added to `docs/MVP-PRD.md`:
### 13.1 Platform requirements (REQ-PLT-XX)
| ID | Requirement |
|---|---|
| REQ-PLT-05 | All processes run as standard user (no sudo) — *already exists* |
| REQ-PLT-10 | The pipeline process shall not exceed 500MB of RSS memory (excluding host-level Ollama) |
| REQ-PLT-15 | The system shall support deployment on a VPS with 2GB total RAM (pipeline + Ollama + OS) |
| REQ-PLT-20 | Ollama inference shall run as a host-level service, not inside the pipeline container |
| REQ-PLT-25 | The pipeline shall use SQLite as the sole database (no PostgreSQL, no Elasticsearch, no Redis) |
### 13.2 Reliability requirements (REQ-REL-XX)
| ID | Requirement |
|---|---|
| REQ-REL-05 | The system operates autonomously without human interaction — *already exists* |
| REQ-REL-10 | Failed source fetches retry with exponential backoff — *already exists* |
| REQ-REL-15 | Previously stored data is not lost on restart or failure — *already exists* |
| REQ-REL-20 | The pipeline completes successfully even if 1+ sources are unavailable — *already exists* |
| REQ-REL-25 | If Ollama is unreachable, ingestion continues and summarization defers to the next run |
| REQ-REL-30 | The pipeline uses flock/PID file to prevent overlapping runs |
### 13.3 Observability requirements (REQ-DIAG-XX)
| ID | Requirement |
|---|---|
| REQ-DIAG-05 | Structured JSON logs with timestamps and severity — *already exists* |
| REQ-DIAG-10 | Health endpoint reports system status and last successful run — *already exists* |
| REQ-DIAG-15 | Per-source success/failure and fetch counts logged per run — *already exists* |
| REQ-DIAG-20 | Webhook alert (Discord/Slack) fires when an adapter fails for 2+ consecutive days |
| REQ-DIAG-25 | run_log table captures per-run metrics queryable via SQL |
### 13.4 Integration requirements (REQ-INT-XX)
| ID | Requirement |
|---|---|
| REQ-INT-05 | Structured, machine-readable output (JSON) consumable by external tools — *already exists* |
| REQ-INT-10 | Adapter layer supports adding new sources without modifying core pipeline logic — *already exists* |
| REQ-INT-15 | REST API exposes GET /trends, /entries, /verdicts endpoints |
| REQ-INT-20 | Daily file drop at configurable path with structured output (YAML + JSON) |
| REQ-INT-25 | MCP server exposes oracle_search, oracle_trends, oracle_verdicts tools (Phase 6) |
### 13.5 Data requirements (REQ-DATA-XX) *(new section)*
| ID | Requirement |
|---|---|
| REQ-DATA-05 | Entries are deduplicated by (source, source_id) with cross-source URL hash deduplication |
| REQ-DATA-10 | FTS5 full-text index on title and extracted_text for keyword search |
| REQ-DATA-15 | Convergence detection: theme appears in ≥3 independent sources within 24h window |
| REQ-DATA-20 | Falsification uses exponential decay scoring (configurable λ), not fixed day thresholds |
| REQ-DATA-25 | Data retention: raw entries 90 days, summaries/convergence 365 days, periodic VACUUM |
### 13.6 Security requirements (REQ-SEC-XX)
| ID | Requirement |
|---|---|
| REQ-SEC-05 | All outbound HTTP uses TLS — *already exists* |
| REQ-SEC-10 | No secrets hardcoded or stored in plaintext — *already exists* |
| REQ-SEC-15 | Least-privilege access for outbound API calls — *already exists* |
| REQ-SEC-20 | SQLite database file permissions set to 600 (owner-only read/write) |
| REQ-SEC-25 | All SQL queries use parameterized statements (no string concatenation) |
### 13.7 Scheduling requirements (REQ-SCH-XX) *(new section)*
| ID | Requirement |
|---|---|
| REQ-SCH-05 | Pipeline entry point is oracle-pipeline.sh (idempotent, single command) |
| REQ-SCH-10 | Default schedule: 13:00 UTC daily |
| REQ-SCH-15 | Exit codes: 0 = success, 1 = partial failure, 2 = total failure |
| REQ-SCH-20 | Overlapping run prevention via flock or PID file check |
---
## 14. Design Decisions Summary (Why)
| Decision | Why | Revisit |
|---|---|---|
| **SQLite over PostgreSQL** | Zero external dependency, single file, FTS5 built-in, handles 1M rows fine. pgvector is premature at this scale. | P6 (if scale demands pgvector) |
| **Ollama host-level** | 150MB container cannot fit Ollama + model (~2GB). Host-level lets pipeline stay within budget. | P2 (if inference needs change) |
| **Flask over FastAPI** | ~1MB vs ~100MB runtime overhead. FastAPI is Phase 2 target; Flask suffices for internal REST API. | P2 (when FastAPI becomes viable) |
| **Keyword co-occurrence (Phase 1)** | Zero-dependency, explainable, works at 150MB. Embeddings (Phase 2) add semantic convergence. | P2 (when embeddings ready) |
| **Fixed themes + catch-all** | BERTopic requires 4GB RAM. Fixed themes with "other" bucket is the pragmatic constraint choice. | P6 (when auto-discovery needed) |
| **Cron over APScheduler (Phase 1)** | OS-level, zero process memory cost. APScheduler is Phase 2 for dynamic scheduling. | P2 (if per-source intervals needed) |
| **HTTP-only adapters** | All 6 sources have programmatic APIs. Playwright adds 300MB+ overhead and fragility. | N/A (stable) |
| **Exponential decay over 7-day rule** | One-line formula, no fixed threshold. Handles fast-dying and slow-burn trends naturally. | N/A (stable) |
| **Deduplication required** | arXiv papers appear on HN/Reddit/Twitter. Without dedup, same signal counted 3× = false convergence. | N/A (stable) |
| **Graceful degradation on Ollama** | Ingestion must not depend on summarization. Store raw data, defer summaries. | N/A (stable) |
## 15. Testing Strategy
### 15.1 Unit Tests
- **Per adapter**: Test each of the 6 adapters with known-good endpoints. Verify:
- Returns valid JSON with required schema fields (source, source_id, title, url, timestamp)
- Handles rate limits gracefully (no infinite loops)
- Correct error codes for 429/503
- **Scoring functions**: Unit tests for exponential decay, convergence scoring, and theme matching. Include edge cases (empty input, negative scores).
### 15.2 Integration Tests
- Run pipeline against seed dataset (100 entries from arXiv + RSS). Verify:
- All adapters fetch successfully
- Deduplication removes duplicates correctly
- Theme detection identifies at least 3 themes
- No critical errors in logs
### 15.3 E2E Tests
- **Cron-to-snapshot**: Run full pipeline via cron, verify output files match expected snapshot (golden file comparison)
- **Stress test**: Run 10 consecutive daily cycles with simulated failures (adapter timeout, Ollama down). Verify graceful degradation and recovery
**Test coverage goal**: 80% of critical paths covered by automated tests. Manual testing for theme quality and summary accuracy.
## 16. Deployment and CI/CD
### 16.1 CI Pipeline
On every push to `MVP-milestone`:
- **Lint**: `ruff check`, `mdlint docs/`
- **Test**: Run unit tests (`pytest tests/`), integration tests with seed dataset
- **Build**: Create Docker image, tag with commit SHA
- **Scan**: Run `pip audit`; fail if critical vulnerabilities (>CVSS 7.0) found
### 16.2 Deployment
- **Local dev**: `docker-compose up` (pipeline container + host-level Ollama)
- **Production**: `docker-compose up -d` + systemd services for inference (`qwythos-gpu0.service`)
- Ollama runs host-level (not in container) due to ~2GB memory requirement
### 16.3 Rollback Procedure
If deployment fails or quality degrades:
1. Stop services: `systemctl stop oracle-pipeline qwythos-gpu0`
2. Restore previous code: `git checkout <good-tag>`
3. Rebuild Docker image from restored code
4. Restart services: `systemctl start oracle-pipeline qwythos-gpu0`
5. Verify health: `curl http://localhost:8081/v1/health`
**Rollback window**: Must complete within 5 minutes of failure detection.
**Tagging**: Each deploy is tagged with semantic versioning (`v1.0.0`, `v1.1.0`) for easy rollback reference.
## 17. Operational Runbooks
### 17.1 Daily Pipeline Verification
At 13:00 UTC after each run:
1. Check logs: `journalctl -u oracle-pipeline --since "today" | grep ERROR`
2. Verify exit code in `/var/log/oracle-pipeline/run_log.txt` — should be 0
3. Confirm output: `ls -lh /output/$(date +%Y-%m-%d)/` — should have `summary.json` and `metrics.json`
4. If any check fails, investigate with the relevant runbook below
### 17.2 Database Recovery from Corruption
If `sqlite3 oracle.db 'PRAGMA integrity_check'` returns errors:
1. Stop pipeline: `systemctl stop oracle-pipeline`
2. Restore from last backup: `cp /backup/oracle.db.bak-YYYYMMDD oracle.db`
3. Verify integrity: `sqlite3 oracle.db 'PRAGMA integrity_check'` — should return "ok"
4. Restart pipeline: `systemctl start oracle-pipeline`
**Prevention**: Daily compressed backups to `/backup/`, retention 30 days.
### 17.3 Adapter Failure Investigation
If adapter fails repeatedly (>2 consecutive runs):
1. Check logs: `journalctl -u oracle-pipeline --since "today" | grep -A5 "Adapter"`
2. Test endpoint manually: `curl -X GET <adapter_url>` — verify HTTP status
3. Check rate limit headers: `curl -I <adapter_url> | grep -i 'x-ratelimit'`
4. If rate-limited: Wait `retry_after` seconds, pipeline retries next cycle
5. Escalate to project lead if >3 consecutive failures
### 17.4 Ollama Service Restart
If Ollama becomes unresponsive:
1. Check status: `systemctl status qwythos-gpu0.service`
2. View logs: `journalctl -u qwythos-gpu0.service --since "today"`
3. Restart service: `systemctl restart qwythos-gpu0.service`
4. Verify health: `curl http://localhost:8081/v1/health` — should return `{ "status": "ready" }`
5. Check GPU: `nvidia-smi` — ensure service actually loaded
**Escalation**: If issue persists after restart, notify project lead.
**Runbook maintenance**: Update runbooks when features change. Document changes in `runbook_changes.md`.
## 18. Risk Register and Assumptions
### 18.1 Risk Register
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Source API changes (rate limits, endpoint shifts) | High | High | Auto-retry with exponential backoff; monitor adapter health; log API changes for review |
| Model quality drift (summarization accuracy degrades) | Medium | Medium | Weekly golden sample comparison; rollback if hallucination rate >10% or summary <50 tokens |
| Disk space exhaustion (backups, logs) | High | High | Automated 30-day retention; alert at 85% usage; compress old backups |
| Single operator bottleneck (manual theme review) | Medium | Medium | Documented theme governance; quarterly review cycle; catch-all "other" bucket |
| Network outage (all adapters fail) | Low | High | Graceful degradation: store raw data without summaries; resume when network returns |
| SQLite corruption during schema migration | Low | Critical | Pre-migration backup; atomic migration scripts; integrity check after each migration |
### 18.2 Key Assumptions
- All 6 source APIs remain stable for at least 6 months (no breaking changes)
- GPU memory available for Qwythos inference (~15GB free on GPU0)
- Network connectivity to all sources is available during pipeline runs
- Single operator can complete manual theme review within 24 hours
- Disk space sufficient for 30-day backup retention (~10GB)
**Risk review**: Quarterly review of this register; update mitigations when new risks emerge.
**Assumption tracking**: If an assumption proves false, document the deviation in `assumption_deviations.md` and reassess risks.
## 19. Dependencies and Tooling
### 19.1 Python Runtime Dependencies
All dependencies pinned to exact versions in `requirements.txt`:
- **requests**: HTTP client for all source APIs
- **beautifulsoup4**: HTML parsing for RSS feeds
- **feedparser**: RSS/Atom feed handling (primary)
- **tenacity**: Retry logic with exponential backoff (used by all adapters)
- **flask**: Internal REST API for Phase 4+ (served on port 8081)
No external database dependencies — SQLite is built into Python.
### 19.2 Runtime System Dependencies
- **Ollama**: Host-level inference service (~2GB RAM, GPU acceleration)
- **systemd**: Service management for pipeline and inference (`oracle-pipeline.service`, `qwythos-gpu0.service`)
- **cron**: Daily schedule trigger (Phase 1); systemd timers for Phase 2
- **sqlite3**: Database CLI for backup/recovery commands
### 19.3 Development Tooling
- **ruff**: Linting and formatting (Python)
- **pytest**: Unit and integration test framework
- **docker**: Containerization for pipeline
- **pip audit / safety**: Dependency vulnerability scanning (CI checks)
### 19.4 Dependency Update Policy
- Critical security updates: Apply within 7 days of CVE disclosure
- Minor version updates: Test in staging before production deployment
- Major version upgrades: Require migration scripts and rollback plan
**Tooling policy**: No new dependencies without approval from project lead; document rationale in `dependency_justifications.md`.
**Lock file**: `requirements.txt` pinned to exact versions for reproducible builds.
---
*Document prepared via multi-model analysis: Qwythos-9B (architectural critique), Qwen3.5-9B (implementation evaluation), and cross-review synthesis. 4 delegations, 2 rounds of debate. Raw reviews saved in the same directory.*
+85
View File
@@ -0,0 +1,85 @@
# Dev Design Review
## Overall Assessment
The current Dev-Design.md is 75-80% ready for delegation. Strong architecture and phased approach, but missing key operational sections.
## Chapter-by-Chapter Feedback
### 1. North Star
**Strengths**: Clear vision.
**Gaps**: No measurable success criteria for MVP.
**Recommendation**: Add success metrics (e.g., Bob spends <15 min/week on research).
### 2. System Architecture
**Strengths**: Good diagram and memory budget.
**Gaps**: No data flow description, error propagation, or external dependency diagram.
**Recommendation**: Add numbered daily run flow and failure modes section.
### 3. Data Layer
**Strengths**: Strong SQLite rationale and retention policy.
**Gaps**: No backup strategy or migration process.
**Recommendation**: Add SQLite backup/restore and schema migration approach.
### 4. Adapter Layer
**Strengths**: Clear HTTP-only decision.
**Gaps**: No formal adapter interface contract.
**Recommendation**: Define minimal Adapter Interface (methods, exceptions).
### 5. Theme Tagging
**Strengths**: Good phased approach.
**Gaps**: No theme governance process.
**Recommendation**: Add subsection on how themes are proposed and maintained.
### 6. Falsification Engine
**Strengths**: Exponential decay is a smart improvement.
**Gaps**: No calibration/tuning process.
**Recommendation**: Add note on how decay parameters are validated.
### 7. Output Layer
**Strengths**: Good consumer separation.
**Gaps**: MCP tools are too high-level; no auth model.
**Recommendation**: Define minimum MCP tools and basic API auth.
### 8. Observability
**Strengths**: Decent start.
**Gaps**: No log retention, alerting thresholds, or SLIs.
**Recommendation**: Add log retention policy and basic alerting matrix.
### 9. Scheduling
**Strengths**: Good comparison.
**Gaps**: No explicit recommendation for Phase 1 vs future.
**Recommendation**: State clear recommendation (cron for Phase 1).
### 10. Inference
**Strengths**: Clear host-level decision.
**Gaps**: No model upgrade/rollback process.
**Recommendation**: Add model upgrade guidance.
### 11. Security
**Strengths**: Basic coverage.
**Gaps**: Container hardening and dependency scanning.
**Recommendation**: Add MVP security baseline subsection.
### 12. Implementation Phases
**Strengths**: Strong.
**Gaps**: No Definition of Done per phase.
**Recommendation**: Add DoD checklist for each phase.
### 13. Backport to PRD
**Strengths**: Useful.
**Recommendation**: Consider moving actual requirement text to PRD to avoid duplication.
### 14. Design Decisions
**Strengths**: Good.
**Recommendation**: Add "Revisit in Phase X" column for key decisions.
## Missing Sections to Add
1. **Testing Strategy** (unit, integration, E2E)
2. **Deployment & CI/CD**
3. **Operational Runbooks**
4. **Risk Register & Assumptions**
5. **Dependencies & Tooling**
## Priority for Next Revision
Focus on adding Testing Strategy, Runbooks, and DoD per phase first. This will make the document truly delegation-ready.
+161
View File
@@ -0,0 +1,161 @@
## Chapter 1: Vision and Scope
### Elevator Pitch
Athena is an autonomous research intelligence engine that cuts through high-volume, fragmented signals by ingesting from multiple sources, surfacing cross-source convergence, and using falsification to distinguish real momentum from noise. While the initial focus is on AI signals, the system is designed to work with any class of signals. It delivers actionable insight into emerging trends and capability gaps while remaining model-agnostic and lightweight enough to run autonomously.
### 1.1 Vision
#### Why are we building it?
The AI space produces an overwhelming volume of new research, tools, discussions, and model releases every day. Individual sources only provide partial views, making it difficult to distinguish genuine, sustained trends from one-day spikes. Without a system that can detect convergence across sources and validate momentum over time, real opportunities tied to emerging capability gaps are missed.
#### What happens if we dont build it?
Without this capability, builders and researchers will continue to operate with fragmented, noisy signals. Early indicators of meaningful trends will remain hidden, decisions will stay reactive, and the ability to spot validated cross-source momentum before it becomes obvious will be lost.
#### When must it be done?
The foundational ability to reliably ingest, score, and validate signals through falsification must be established before meaningful trend detection and opportunity mapping can occur. This forms the core of the MVP and must be in place to enable the system to deliver on its intended value.
### 1.2 Personas and Archetypes
See committed document:
**`docs/Personas-and-Archetypes.md`** (on `MVP-milestone` branch)
**Summary of scoped personas and archetypes for MVP:**
**Personas**
- Pers-1 (Bob) Sector Trend Tracker (New to AI)
- Pers-2 (Alice) Content Creator
- Pers-3 (Sam) Hermes Research Agent
**Archetypes**
- Arch-1 (Small Scrappy VPS)
- Arch-2 (Research Consumption Layer)
All user stories in this PRD are scoped to combinations of the above.
### 1.3 Use Case Priority Taxonomy
This PRD focuses on defining the core functionality required for MVP. It also catalogs use cases and requirements across V1.0 V1.5 to maintain context. The primary goal is to deliver a working MVP, with future PRDs derived from the remaining prioritized content.
We will use the following prioritization model:
- **MVP**: The short list of P1 use cases required to prove the concept with a working prototype.
- **P1**: Use cases that are fundamental to successfully implementing the product vision.
- **P2**: Use cases that add strength, convenience, and quality to the product vision.
- **P3**: Use cases that bring additional value but can be cut if time or resource constrained.
## Chapter 2: User Stories (Bob)
These user stories are based on the personas and archetypes document contained in this repo.
Chapter 2.1 - Bob's user stories
**As Bob, I want to…**
**Bob-1.** Automatically receive daily updates on new AI innovations without having to manually check multiple sources.
**Bob-5.** See emerging trends and differentiate durable signal from temporary or artificial hype.
**Bob-10.** See when the same idea or pattern is appearing across multiple independent sources (GitHub, arXiv, Reddit, HN, HF).
**Bob-15.** Identify emerging capability gaps or opportunities early, before they become widely obvious.
**Bob-20.** Have research that gives me confidence it is exhaustive and vetted.
**Bob-25.** Adjust or alter the underlying data feeds and weights so I can tune the accuracy and relevance of the output.
**Bob-30.** Understand why a particular signal is considered strong or weak (e.g., cross-source convergence or falsification results).
Chapter 2.2 - Alice's user stories
As Alice, I want to…
Alice-1. Integrate deep, vetted research directly into my existing content production pipeline so I can reduce manual research time.
Alice-5. Query the research system with follow-up questions to explore specific angles or topics on demand.
Alice-10. Have my tools automatically receive curated, high-signal research so I can focus on content creation instead of information filtering.
Alice-15. Get research outputs in a structured format that my existing AI tools and workflows can consume without manual reformatting.
Alice-20. Quickly surface non-obvious insights and patterns from research data to develop more compelling content angles.
Alice-25. Control which research sources and signals are prioritized so the output stays aligned with my content focus and audience.
Alice-30. Understand the reasoning and supporting evidence behind key research findings so I can speak to them confidently in my content.
## Chapter 3: Requirements
Requirements defined as what the product / system must do, differentiated from what the persona can accomplish. Requirements are defined to meet the needs of use cases as well as the architectural system design.
High level design (refer to ***TBD_Design.MD for full design details)
High-Level Design
.
├── Runtime Environment
│ ├── Linux
│ └── Docker (containerized)
├── Core Components
│ ├── Database: SQLite
│ ├── Scheduling: Cron
│ └── Runtime: Python
├── Connectivity
│ ├── Outbound (Internet)
│ │ ├── HTTP client for data feeds (RSS, cURL, optional Playwright)
│ │ └── OpenAI-compatible inference endpoints
│ ├── Inbound (Internet)
│ │ └── HTTP server endpoint (MCP + external consumers)
│ └── Internal (Intranet)
│ └── HTTP client for local inference (e.g. Hermes)
├── Storage
│ ├── File system (daily digest artifacts stored outside container)
│ └── Temporary working storage during pipeline execution
├── Configuration & Secrets
│ ├── Research topic manifest (feeds, URLs, declarations)
│ ├── System settings (YAML)
│ └── Secrets (.env)
├── Business Logic / Pipeline Flow
│ ├── Starting trigger
│ ├── Preflight checks
│ ├── Query feeds → Temporary result storage
│ ├── Vet and promote final results to database
│ ├── Optional daily digest generation
│ └── Cleanup and sleep
└── Observability
├── Structured logging
├── Diagnostics and instrumentation (inside Docker)
└── Health/status reporting
3.1 Setup and configuration (REQ-SNC-XX)
Requirements for initial setup, deployment configs, updating, and uninstall
REQ-SNC-05 -, with outbound access to the internet and in/outbound access to the underlying OS network
REQ-SNC-10 - The installation process shall be a single command which can be run interactively or silently
REQ-SNC-15 - The insallation shall utilize best-practice settings and secrets storage
REQ-SNC-20 -
3.2 Platform requirements (REQ-PLT-XX)
REQ-PLT-05 - All processes will run as standard user (no admin / sudo elevation necessary)
REQ-PLT-10 - ...
REQ-PLT-15 - The system shall be Docker based limited to 150MB of memory
Requirements addressing what OS and hardware support is in scope
3.3 Performance and scalability (REQ-PERF-XX)
3.4 Instrumenation and diagnostics (REQ-DIAG-XX)
3.5
### 3.1 Reliability
REQ-REL-05: Once setup and configured, the system will reliably operate without interaction from the user.
REQ-REL-10: The system shall automatically retry failed source fetches with exponential backoff.
REQ-REL-15: The system shall not lose previously stored data on restart or failure.
REQ-REL-20: The daily pipeline shall complete successfully even if one or more sources are unavailable.
3.3 Observability and Diagnostics
REQ-DIAG-05: The system shall produce structured logs with timestamps and severity levels.
REQ-DIAG-10: A health check endpoint or command shall report overall system status and last successful run.
REQ-DIAG-15: Run logs shall capture per-source success/failure and basic metrics (items fetched, stored, failed).
3.4 Security
REQ-SEC-05: All external HTTP calls shall use TLS.
REQ-SEC-10: No secrets shall be hardcoded or stored in plaintext.
REQ-SEC-15: The system shall support least-privilege access for outbound API calls.
3.5 Integration and extensibility
REQ-INT-05: The system shall expose research output in a structured, machine-readable format (e.g., JSON files or API) consumable by external tools.
REQ-INT-10: The adapter layer shall support adding new sources without modifying core pipeline logic.
+59
View File
@@ -0,0 +1,59 @@
# Personas and Archetypes (Derived from /main)
**Status**: First Draft Derived from existing docs on `main`
**Source**: README.md + whitepaper.md (branch: main)
**Date**: 2026-07-08
## Overview
This document extracts the implied users and operating contexts directly from the current documentation on the `main` branch. It serves as the baseline before we expand or refine.
---
## Personas
### Pers-1 (Bob) Sector Trend Tracker (New to AI)
- Is relatively new to AI and the broader space.
- Wants to stay current with trends in AI (and potentially other sectors) without getting overwhelmed.
- Needs a way to keep up with the high volume of new research, tools, and discussions with minimal ongoing effort.
- Benefits from a system that filters noise and surfaces what actually matters.
### Pers-2 (Alice) Content Creator
- Runs a YouTube channel and an X account with 25k followers.
- Goal is to grow her audience significantly (targeting 1M followers).
- Needs help doing research across AI and related topics.
- Wants to convert research signals into interesting, timely, and compelling content for her audience.
### Pers-3 (Sam) Hermes Research Agent
- Is a Hermes agent profile with its own memory and endpoint connection.
- Acts as the dedicated research team member for an AI-first development team.
- Needs to stay current on a defined market segment (AI for the MVP; extensible to other segments later).
- Consumes structured signals from Athena to support ongoing research and decision-making within the team.
---
## Archetypes (Operating Environments)
### Arch-1 (Small Scrappy VPS)
- Small, low-budget, and scrappy VPS environment.
- Used primarily for learning and early prototyping.
- Requires a small-footprint workload that can be memory-constrained so it doesnt destabilize the host system.
### Arch-2 (Research Consumption Layer)
- Functions as a consumption layer for Athenas research output.
- Designed to support downstream AI systems (examples: MCP tools, LoRA adapters, or other agent profiles).
- Focuses on making Athenas signals and summaries easily consumable by other systems rather than direct human use.
---
## Notes & Limitations (from /main)
- The current documentation does **not** describe team or multi-user usage.
- Emphasis is on autonomous operation and signal integrity.
- Polished human-facing interfaces (e.g., daily digest) are not yet built.
---
## Next Steps
This version incorporates the updated personas and archetypes.
+42
View File
@@ -0,0 +1,42 @@
# Prioritized Task List — Athena MVP
Tied to the personas and requirements in `MVP-PRD.md`. Ordered by phase; within each phase, roughly in the order they should be tackled.
## Phase 1: Get Running Daily
- [ ] Verify the cron entry (`oracle-pipeline.sh`) fires reliably at 13:00 UTC under Hermes
- [ ] Confirm `pipeline.py` runs the full ingest → store → summarize → score cycle without manual intervention
- [ ] Add a lock/guard so a slow run can't overlap with the next day's cron trigger
- [ ] Validate all 6 adapters (arxiv, github, huggingface, hackernews, reddit, rss_feeds) independently — one adapter failing shouldn't kill the whole run
- [ ] Confirm environment-only secrets (`GITHUB_TOKEN`, `HUGGINGFACE_TOKEN`) resolve correctly in the cron context (cron environments are often stripped down compared to an interactive shell)
## Phase 2: Core Functionality
- [ ] Confirm `schema.sql` initializes `oracle.db` cleanly and stays idempotent across repeated runs
- [ ] Verify `theme_scan.py`'s 4-theme tagging (tool-call, context, compute, trust) against a few real days of data
- [ ] Confirm the falsification counter (new arrivals per cycle) is genuinely idempotent — re-running against unchanged data must yield 0 new
- [ ] Wire `summarize.py` to degrade gracefully when the Ollama endpoint (`llama3.2:1b`) isn't reachable — ingestion, scoring, and theme-scan must keep running without it
- [ ] Confirm `archive.py`'s cold-storage rotation doesn't delete data still needed inside the 7-day falsification window
## Phase 3: Observability & Reliability
- [ ] Add structured logging per pipeline stage (ingest, store, summarize, score) with pass/fail per adapter
- [ ] Surface theme-scan counts (new arrivals per theme per cycle) somewhere inspectable, not just buried in log files
- [ ] Add a daily heartbeat/health check so a silent failure (e.g. cron didn't fire at all) is detectable rather than just showing up as missing data later
- [ ] Decide and implement retry/backoff behavior for adapters that hit rate limits (especially GitHub without a token: 60/hr)
## Phase 4: Human Consumption Layer
- [ ] Extend `query.py` to support Bob's cross-source convergence lookups and Alice's curated-research pulls
- [ ] Define the output format(s) for a "trend confirmed" vs. "trend killed" verdict (per the 7-day dead-thesis rule)
- [ ] Decide how Alice's content pipeline actually consumes Athena's output — file drop, API, direct DB read — this is currently undefined
## Phase 5: Validation & UAT
- [ ] Run the pipeline unattended for at least one full 7-day falsification window
- [ ] Manually verify at least one theme through to a real "confirmed" or "killed" verdict
- [ ] Walk Bob's and Alice's user stories from `MVP-PRD.md` end-to-end against real output, not synthetic data
- [ ] Confirm memory stays under the 150MB cap under real daily load, not just in a light dev test
## Phase 6: Hermes Integration
- [ ] Confirm `oracle-pipeline.sh`'s contract matches what Hermes cron expects (exit codes, output location)
- [ ] Decide how Hermes is notified on pipeline failure vs. success — not yet specified
- [ ] Confirm the non-root execution requirement is actually satisfied inside the Hermes-invoked environment, not just in local Docker testing
---
*Draft prepared by Claude from the MVP-PRD, Personas doc, and README/whitepaper on `main`. Open items flagged "not yet specified" need a decision before Phase 46 can be considered done.*