# Athena — AI Research Intelligence Engine > Multi-source research ingestion, pattern detection, and hypothesis falsification pipeline. > Autonomous daily operation: ingest → summarize → theme-scan → flag weak signals. ## Repo - **Location:** `Tony_tech/athena-oracle` (public, Gitea) - **URL:** http://localhost:3000/Tony_tech/athena-oracle - **Branch:** `main` - **Origin:** mirrors `~/oracle` (local working copy) ## What it does Athena runs on a daily cron (13:00 UTC) and continuously ingests from 6 sources, then applies a signal-scoring + falsification loop to surface real AI research momentum rather than source-expansion noise. | Component | File | Purpose | |-----------|------|---------| | Pipeline | `pipeline.py` | Orchestrates ingest → store → summarize → score | | Adapters | `adapters/` | arxiv, github, huggingface, hackernews, reddit, rss_feeds | | Theme scan | `theme_scan.py` | Cross-source trend detection + idempotent falsification | | Query | `query.py` | Interactive lookup against the store | | Archive | `archive.py` | Cold-storage rotation | | Summarize | `summarize.py` | Summarization via any available inference model | | Schema | `schema.sql` | SQLite store definition | | Cron entry | `oracle-pipeline.sh` | Wrapper invoked by Hermes cron | ## Inference model strategy Athena is **model-agnostic** — it uses whatever inference backend is available at run time, whether free or paid. There is no hard dependency on a single provider. `summarize.py` currently targets a local Ollama endpoint (`llama3.2:1b`) when present. The pipeline is designed so the summarization backend can be swapped for any model we can reach — local GPU, a paid API, or a free-tier endpoint — without changing the ingestion, scoring, or theme-scan logic. When no inference backend is reachable, the summarization step is skipped; ingestion, scoring, and theme-scan continue uninterrupted. To wire in a different backend, implement the same `summarize(text) -> (summary, model)` contract that `summarize_with_ollama` satisfies, and add the dispatch in `process_card`. ## Data handling - `oracle.db`, `logs/`, `.env`, `__pycache__/` are **git-ignored** (not committed). - API tokens (`GITHUB_TOKEN`, `HUGGINGFACE_TOKEN`) are read from environment only — never hardcoded. ## Setup ```bash pip install -r requirements.txt # if present; else deps are stdlib + requests export GITHUB_TOKEN=... # optional, raises rate limit 60→5000/hr python3 pipeline.py # manual run ``` ## Architecture detail See `whitepaper.md` for full system design, scoring methodology, and the verification discipline that keeps adapters honest.