refactor(adapters): centralize curation in config/queries.json (issue #7)

- adapters/__init__.py: add load_queries() + source_config() (stdlib json,
  safe fallback to {} on missing/corrupt config so pipeline never crashes).
- config/queries.json: per-adapter blocks (hackernews.keywords, arxiv.categories,
  reddit.subreddits, rss.feeds+keywords, github.search_terms). JSON (not
  yaml) to honor Athena's dependency-free runtime; PyYAML avoided.
- hackernews: drop class AI_KEYWORDS + the DUPLICATE inline list inside
  _is_ai_relevant() (the internal drift Ty flagged). Now loads self.ai_keywords
  from config. Simplified matching to single substring pass (boundary variants
  'ai ',' ai','ai-','-ai' approximate word-boundary; dropped the niche
  'compute+tech-context' guard as not worth centralizing).
- reddit: DEFAULT_SUBREDDITS kept as fallback; __init__ prefers config.
- arxiv: DEFAULT_CATEGORIES kept as fallback; __init__ prefers config.
- rss: FEEDS + AI_KEYWORDS kept as module fallbacks; __init__ prefers config.
  Keywords stay regex form (re.search) as in original.
- github: trending queries moved to config search_terms; fallback retained.
- DELETE reddit_proof.py: standalone PoC v5 at repo root, own main()+init_db()
  + direct INSERT OR REPLACE, NOT in cron, NOT imported anywhere -> dead
  code. Also removes its byte-duplicate SUBREDDITS.

NOTE: fallback class constants remain intentionally (issue #7 cut #5: safe
rollout). Curation VALUES now live in one file; the constants are inert
unless config/queries.json is missing.

Verified: all adapters compile; config loads (HN 46 kw, RSS 10 feeds);
full dry-run fetches all 6 sources; grep confirms HN internal dup list gone.
This commit is contained in:
Epictetus
2026-07-10 17:42:06 +00:00
parent 7ee1af3d7b
commit a9fbe27178
8 changed files with 115 additions and 386 deletions
+49
View File
@@ -0,0 +1,49 @@
{
"sources": {
"hackernews": {
"keywords": [
"language model", "deep learning", "foundation model", "retrieval augmented",
"code generation", "context length", "context window", "attention mechanism",
"llm", "gpt-", "gpt ", "rag ", "rag.", "vlm", "vla",
"openai", "anthropic", "deepseek", "meta ai", "xai", "ponytail",
"inference", "transformer", "diffusion", "alignment", "fine-tun",
"embed", "pretrain", "post-train", "multimodal", "reasoning",
"ai ", " ai", "ai-", "-ai",
"agent", "agents", "neural", "autonomous",
"compute", "training run", "computer use", "coding agent",
"local-llm", "local llama", "llama "
]
},
"arxiv": {
"categories": ["cs.AI", "cs.LG", "cs.CL"]
},
"reddit": {
"subreddits": ["MachineLearning", "artificial", "LocalLLaMA", "Startups"]
},
"rss": {
"feeds": [
["rss:techcrunch", "TechCrunch AI", "https://techcrunch.com/category/artificial-intelligence/feed/"],
["rss:venturebeat", "VentureBeat AI", "https://venturebeat.com/category/ai/feed/"],
["rss:theverge", "The Verge AI", "https://www.theverge.com/rss/ai-artificial-intelligence/index.xml"],
["rss:ainews", "AI News", "https://www.artificialintelligence-news.com/feed/"],
["rss:decoder", "The Decoder", "https://www.the-decoder.com/feed/"],
["rss:mittr", "MIT Tech Review AI", "https://www.technologyreview.com/topic/artificial-intelligence/feed/"],
["rss:openai", "OpenAI Blog", "https://openai.com/blog/rss.xml"],
["rss:anthropic", "Anthropic News", "https://www.anthropic.com/rss/news.xml"],
["rss:googleai", "Google AI Blog", "https://blog.google/technology/rss.xml"],
["rss:metaai", "Meta AI Blog", "https://ai.meta.com/blog/rss.xml"]
],
"keywords": [
"\\bai\\b", "\\bmachine learning\\b", "\\bdeep learning\\b", "\\bneural\\b",
"\\bgenerative ai\\b", "\\bgenerative\\b", "\\bllm\\b", "\\blarge language\\b",
"\\bfoundation model\\b", "\\btransformer\\b", "\\baugmented\\b",
"\\bagent\\b", "\\bautonomous\\b", "\\bmcp\\b", "\\bfunction call\\b",
"\\btool use\\b", "\\brai\\b", "\\bretrieval\\b",
"\\binference\\b", "\\bmodel\\b", "\\bembedding\\b", "\\btoken\\b"
]
},
"github": {
"search_terms": ["machine-learning", "deep-learning", "llm", "ai-agent", "transformer"]
}
}
}