fix(adapters): shared retry helper + run_log failure_class + enable RSS (issues #1 #2 #9)

- adapters/__init__.py: add http_get() unified retry (429/5xx only, max 2
  attempts, capped exp backoff) + AdapterHTTPError carrying failure_class;
  SourceAdapter.last_failure_class set on failure for pipeline capture.
- arxiv/github/huggingface/hackernews/reddit: route HTTP through http_get.
  Preserves GitHub 403 rate-limit retry and Reddit 403/429 fast-bail.
- schema.sql + pipeline.py: add run_log.failure_class column; rollup most-
  severe class across sources (5xx>4xx>429>error>zero_fetch>ok).
- pipeline.py: ENABLE RSS in ENABLED_SOURCES (was registered, disabled).
- RSS smoke test surfaced 3 broken feeds (anthropic 404, googleai 404,
  metaai 301) — left as-is, captured in feed_failures; URL fix is separate
  discovery task, not guessed.

Verified: full dry-run fetches all 6 sources; github live fetch OK;
Reddit 429 fast-bail preserved; no import/syntax errors.
This commit is contained in:
Epictetus
2026-07-10 16:34:05 +00:00
parent 7ee1af3d7b
commit 23cce4d609
8 changed files with 209 additions and 137 deletions
+13 -18
View File
@@ -34,7 +34,7 @@ import urllib.request
import urllib.error
from datetime import datetime, timedelta, timezone
from adapters import SourceAdapter
from adapters import SourceAdapter, http_get, AdapterHTTPError
class HuggingFaceAdapter(SourceAdapter):
@@ -84,24 +84,19 @@ class HuggingFaceAdapter(SourceAdapter):
return headers
def _request(self, path: str, max_retries: int = 2) -> list | dict | None:
"""Make a GET request to the HF API."""
"""GET via shared retry helper (retries 429/5xx)."""
url = f"{self.BASE}{path}"
req = urllib.request.Request(url, headers=self._headers())
for attempt in range(max_retries + 1):
try:
with urllib.request.urlopen(req, timeout=20) as resp:
return json.loads(resp.read().decode("utf-8"))
except (urllib.error.HTTPError, urllib.error.URLError) as e:
if attempt < max_retries:
time.sleep(3 * (attempt + 1))
continue
print(f" HF API error: {e}")
return None
except Exception as e:
print(f" Request error: {e}")
return None
return None
try:
raw = http_get(url, headers=self._headers(), timeout=20,
max_retries=max_retries, owner=self)
except AdapterHTTPError as e:
print(f" {e.failure_class}: HF {path}")
return None
try:
return json.loads(raw.decode("utf-8"))
except Exception as e:
print(f" HF decode error: {e}")
return None
def _is_ai_relevant(self, model: dict) -> bool:
"""Check if a model/dataset is AI/ML relevant.