07c5f9a5c2
- New oracle/ package (11 modules) with unified CLI (python -m oracle) - Source tiers: Tier 1 (arxiv/github/hf), Tier 2 (rss/hn), Tier 3 (reddit) - Composite verdicts: PUBLISH/WATCH/ARCHIVE/DROP based on signal score + age - Content-hash dedup: SHA-256[:16] normalized, atomic at insert time - Multi-variant editions: 4 YAML configs (default/research/devops/brief) - Variant engine: filter → rank → render (HTML + JSON, themed) - Per-adapter timeout (10s) + threading fallback - Consolidated 12 root scripts → thin wrappers + oracle/ package - Archived stale scripts (_engagement, _live_compare, reddit_proof) - Updated .gitignore, README.md, schema.sql
198 lines
8.4 KiB
Markdown
198 lines
8.4 KiB
Markdown
# Athena Oracle — AI Research Intelligence Pipeline
|
||
|
||
> Multi-source AI news aggregation, scoring, and multi-variant edition rendering.
|
||
> One pipeline → multiple audience-specific editions.
|
||
|
||
**Live site:** https://ai-oracle.com (rendered from this pipeline)
|
||
|
||
## Architecture
|
||
|
||
```
|
||
┌─────────────────────────────────────────────────────┐
|
||
│ adapters/ │
|
||
│ arxiv · github · hackernews · reddit │
|
||
│ huggingface · rss_feeds │
|
||
└───────────────┬─────────────────────────────────────┘
|
||
│ fetch(limit, timeout=10s)
|
||
▼
|
||
┌─────────────────────────────────────────────────────┐
|
||
│ oracle/ — Core Package │
|
||
│ │
|
||
│ scoring.py Pure-rule component scoring (0-10) │
|
||
│ dedup.py Content-hash dedup + verdict engine │
|
||
│ variants.py Multi-variant edition engine │
|
||
│ render.py HTML/JSON variant renderer │
|
||
│ summarize.py Source-aware text summarization │
|
||
│ recency.py Age-based freshness gate │
|
||
│ themes.py Theme-based trend tracking │
|
||
│ archive.py Soft-cap entry archival │
|
||
│ db.py Schema management + migrations │
|
||
│ config.py Centralized configuration │
|
||
│ cli.py Unified CLI (python -m oracle) │
|
||
└───────────────┬─────────────────────────────────────┘
|
||
│
|
||
▼
|
||
┌─────────────────────────────────────────────────────┐
|
||
│ oracle.db — SQLite (784+ entries) │
|
||
│ Columns: signal_score · final_score · content_hash │
|
||
│ · verdict · source_tier · summary │
|
||
└─────────────────────────────────────────────────────┘
|
||
│
|
||
▼
|
||
┌─────────────────────────────────────────────────────┐
|
||
│ variants/ — Edition Configs (YAML) │
|
||
│ │
|
||
│ default.yaml Full feed, clickability-ranked │
|
||
│ research.yaml arXiv + HF papers, signal-ranked │
|
||
│ devops.yaml Shipping tools, 7-day window │
|
||
│ brief.yaml PUBLISH verdict only, top 8 │
|
||
└─────────────────────────────────────────────────────┘
|
||
```
|
||
|
||
## Quick Start
|
||
|
||
```bash
|
||
# Run the ingestion pipeline
|
||
python -m oracle ingest
|
||
|
||
# Generate summaries for pending entries
|
||
python -m oracle summarize
|
||
|
||
# Render all variant editions (preview)
|
||
python -m oracle render --all --dry-run
|
||
|
||
# Render a single variant
|
||
python -m oracle render --variant research
|
||
|
||
# Deploy to production
|
||
python -m oracle render --all --webroot /var/www/html
|
||
|
||
# Check system health
|
||
python -m oracle health
|
||
```
|
||
|
||
## CLI Commands
|
||
|
||
| Command | Description |
|
||
|---------|-------------|
|
||
| `ingest` | Fetch from adapters, score, store to oracle.db |
|
||
| `summarize` | Generate summaries for entries without one |
|
||
| `query top N` | Top N entries by signal score |
|
||
| `query search "text"` | Keyword search across titles/summaries |
|
||
| `query recent --hours 24` | Recent entries |
|
||
| `render` | Render variant editions (HTML + JSON) |
|
||
| `archive` | Soft-cap archival of old entries |
|
||
| `themes` | Theme-based trend tracking |
|
||
| `dedup` | Content-hash dedup + verdict management |
|
||
| `health` | System health check |
|
||
|
||
### Ingest Options
|
||
```bash
|
||
python -m oracle ingest --sources github,arxiv # Specific sources only
|
||
python -m oracle ingest --limit 5 # 5 entries per source
|
||
python -m oracle ingest --dry-run # Fetch but don't store
|
||
```
|
||
|
||
### Render Options
|
||
```bash
|
||
python -m oracle render --list # Show available variants
|
||
python -m oracle render --variant brief # Single variant
|
||
python -m oracle render --all # All variants
|
||
python -m oracle render --all --dry-run # Preview mode
|
||
python -m oracle render --all --webroot /var/www # Production deploy
|
||
```
|
||
|
||
## Scoring Engine
|
||
|
||
Pure-rule component scoring — no embeddings or LLM required.
|
||
|
||
| Component | Weight | Description |
|
||
|-----------|--------|-------------|
|
||
| Shipping | 20% | Code releases, benchmarks, working demos |
|
||
| Utility | 20% | Practical tools, frameworks, integrations |
|
||
| Replication | 25% | Reproducible research, open datasets |
|
||
| Enthusiast | 20% | Community buzz, notable figures |
|
||
| Novelty | 15% | First-of-its-kind, paradigm shifts |
|
||
|
||
**Hype Penalty:** Caps at 45% to prevent buzzwords from dominating.
|
||
|
||
**Signal Score:** 0–10 scale (per-adapter normalization → final composite).
|
||
|
||
## Source Tiers (World Monitor Pattern)
|
||
|
||
| Tier | Sources | Rationale |
|
||
|------|---------|-----------|
|
||
| **Tier 1** (PRIMARY) | arxiv, github, huggingface | Peer-reviewed research, official code releases, model registry |
|
||
| **Tier 2** (SECONDARY) | rss, hackernews | Curated tech media, curated community |
|
||
| **Tier 3** (TERTIARY) | reddit | User-generated discussion |
|
||
|
||
## Composite Verdicts
|
||
|
||
Entries are classified based on signal score + age:
|
||
|
||
| Verdict | Minimum Score | Maximum Age | Meaning |
|
||
|---------|--------------|-------------|---------|
|
||
| **PUBLISH** | ≥ 6.0 | ≤ 48h | High-signal, fresh — front page material |
|
||
| **WATCH** | ≥ 4.0 | ≤ 168h | Solid signal — worth tracking |
|
||
| **ARCHIVE** | ≥ 2.0 | ≤ 720h | Historical value — keep for reference |
|
||
| **DROP** | any | > 720h | Stale — exclude from active feeds |
|
||
|
||
## Content-Hash Dedup
|
||
|
||
SHA-256 content hashing (first 16 hex chars) for cross-source duplicate detection:
|
||
- Normalized whitespace before hashing
|
||
- Applied atomically at ingest time via `adapters/_store.py`
|
||
- Prevents the same story from appearing multiple times across sources
|
||
|
||
## Variant Editions
|
||
|
||
Each variant is a YAML config defining:
|
||
|
||
- **Filters:** verdicts, sources, min_score, max_age_h, max_items
|
||
- **Ranking:** clickability | signal_score | verdict_priority | freshness (with half-life decay)
|
||
- **Display:** theme, accent color, logo, show/hide score/tier/verdict badges
|
||
|
||
Create a new edition by adding a YAML to `variants/` — no code changes needed.
|
||
|
||
## Database
|
||
|
||
SQLite (`oracle.db`) with 784+ entries across 6 sources. Schema includes:
|
||
- Entry metadata (title, url, source, extracted_text, raw_metadata)
|
||
- Scoring (signal_score, component scores, final_score, actionability_score)
|
||
- Dedup/verdict (content_hash, verdict, source_tier)
|
||
- Summarization (summary JSON with one_liner, key_points, implications)
|
||
- Categorization (category_tags, bucket, narrative_id, topic_id)
|
||
|
||
## Sprint Log
|
||
|
||
### Sprint 0 — Foundation (2026-07-22) ✅
|
||
- Package restructure: `oracle/` + `python -m oracle` CLI
|
||
- Per-adapter timeout (10s) + threading fallback
|
||
- Source confidence tiers (3-tier system)
|
||
- Content-hash dedup (SHA-256[:16])
|
||
- Composite verdicts (PUBLISH/WATCH/ARCHIVE/DROP)
|
||
- Wired into pipeline: atomic hash/tier/verdict at insert time
|
||
- Consolidated 12 root scripts → thin wrappers + oracle/ package
|
||
|
||
### Sprint 1 — Multi-variant Editions (2026-07-22) ✅
|
||
- Variant engine: `oracle/variants.py`
|
||
- 4 default editions: default, research, devops, brief
|
||
- Variant-aware HTML/JSON renderer with theme support
|
||
- CLI: `render --variant`, `--all`, `--list`, `--webroot`
|
||
- Verified: all 4 variants render correctly with proper filtering
|
||
|
||
## Dependencies
|
||
|
||
- Python 3.10+
|
||
- PyYAML (`pip install pyyaml`)
|
||
- Trafilatura (for text extraction in adapters)
|
||
- Feedparser (RSS feeds)
|
||
|
||
## Backward Compatibility
|
||
|
||
Old scripts (`pipeline.py`, `summarize.py`, `query.py`, etc.) are thin wrappers that delegate to the unified CLI. Existing cron jobs calling `python3 pipeline.py` continue to work without changes.
|
||
|
||
## License
|
||
|
||
AGPL v3 (see LICENSE for details)
|