Athena — AI Research Intelligence Engine

Multi-source research ingestion, pattern detection, and hypothesis falsification pipeline. Autonomous daily operation: ingest → summarize → theme-scan → flag weak signals.

Repo

What it does

Athena runs on a daily cron (13:00 UTC) and continuously ingests from 6 sources, then applies a signal-scoring + falsification loop to surface real AI research momentum rather than source-expansion noise.

Component File Purpose
Pipeline pipeline.py Orchestrates ingest → store → summarize → score
Adapters adapters/ arxiv, github, huggingface, hackernews, reddit, rss_feeds
Theme scan theme_scan.py Cross-source trend detection + idempotent falsification
Query query.py Interactive lookup against the store
Archive archive.py Cold-storage rotation
Summarize summarize.py Local summarization (Ollama — currently offline)
Schema schema.sql SQLite store definition
Cron entry oracle-pipeline.sh Wrapper invoked by Hermes cron

Local inference status

summarize.py uses Ollama only (DeepSeek fallback removed — OpenRouter 402). Ollama is currently unavailable; summarization steps are skipped until it returns. Pipeline ingestion, scoring, and theme-scan continue without it.

Data handling

  • oracle.db, logs/, .env, __pycache__/ are git-ignored (not committed).
  • API tokens (GITHUB_TOKEN, HUGGINGFACE_TOKEN) are read from environment only — never hardcoded.

Setup

pip install -r requirements.txt   # if present; else deps are stdlib + requests
export GITHUB_TOKEN=...            # optional, raises rate limit 60→5000/hr
python3 pipeline.py                # manual run

Architecture detail

See whitepaper.md for full system design, scoring methodology, and the verification discipline that keeps adapters honest.

S
Description
AI Research Oracle (Athena) — multi-source ingestion, pattern detection, hypothesis falsification pipeline.
Readme 78 MiB
Languages
Python 51.6%
HTML 48%
Shell 0.4%