- docs/vision/VISION_LOCK_V1.md: permanent editorial mission, audience (builders/tinkerers/
local-AI users/home-lab/agent devs), Editorial Test ('what would an enthusiast DO?'),
architectual constraint (editorial vision is primary system), priority order.
- docs/vision/ADR-0007_EDITORIAL_PIVOT.md: accepted pivot record, before/after table,
Sprint 1 evidence (high-signal = actionable, low-signal = spectator content).
- docs/vision/EDITORIAL_GUIDELINES.md: operating rules; score is draft, taste overrides;
hand-before-machine; memory must store 'what can be done' not 'what happened'.
- docs/vision/SECTION_DEFINITIONS.md: 5 sections (What Shipped / Run It Locally /
Benchmarks & Builds / Problem Solved / Worth Trying Tonight).
- docs/vision/SPRINT_1_FINDINGS.md: reference evidence, 45/200 UNCATEGORIZED taxonomy gap.
- Issue_001.md: handcrafted prototype edition from Sprint 1 stories; benchmark for all future automation.
This is the architectural directive. No scoring/narrative/memory work proceeds except in
service of the Vision Lock. Sprint 2 (Lens) stays blocked pending manual review tallies.
Athena — AI Research Intelligence Engine
Multi-source research ingestion, pattern detection, and hypothesis falsification pipeline. Autonomous daily operation: ingest → summarize → theme-scan → flag weak signals.
Repo
- Location:
Tony_tech/athena-oracle(public, Gitea) - URL: http://localhost:3000/Tony_tech/athena-oracle
- Branch:
main - Origin: mirrors
~/oracle(local working copy)
What it does
Athena runs on a daily cron (13:00 UTC) and continuously ingests from 6 sources, then applies a signal-scoring + falsification loop to surface real AI research momentum rather than source-expansion noise.
| Component | File | Purpose |
|---|---|---|
| Pipeline | pipeline.py |
Orchestrates ingest → store → summarize → score |
| Adapters | adapters/ |
arxiv, github, huggingface, hackernews, reddit, rss_feeds |
| Theme scan | theme_scan.py |
Cross-source trend detection + idempotent falsification |
| Query | query.py |
Interactive lookup against the store |
| Archive | archive.py |
Cold-storage rotation |
| Summarize | summarize.py |
Summarization via any available inference model |
| Schema | schema.sql |
SQLite store definition |
| Cron entry | oracle-pipeline.sh |
Wrapper invoked by Hermes cron |
Inference model strategy
Athena is model-agnostic — it uses whatever inference backend is available at run time, whether free or paid. There is no hard dependency on a single provider.
summarize.py currently targets a local Ollama endpoint (llama3.2:1b) when
present. The pipeline is designed so the summarization backend can be swapped for
any model we can reach — local GPU, a paid API, or a free-tier endpoint — without
changing the ingestion, scoring, or theme-scan logic. When no inference backend is
reachable, the summarization step is skipped; ingestion, scoring, and theme-scan
continue uninterrupted.
To wire in a different backend, implement the same summarize(text) -> (summary, model)
contract that summarize_with_ollama satisfies, and add the dispatch in
process_card.
Data handling
oracle.db,logs/,.env,__pycache__/are git-ignored (not committed).- API tokens (
GITHUB_TOKEN,HUGGINGFACE_TOKEN) are read from environment only — never hardcoded.
Setup
pip install -r requirements.txt # if present; else deps are stdlib + requests
export GITHUB_TOKEN=... # optional, raises rate limit 60→5000/hr
python3 pipeline.py # manual run
Architecture detail
See whitepaper.md for full system design, scoring methodology, and the
verification discipline that keeps adapters honest.