Compare commits
1 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 28b0c483b8 |
@@ -0,0 +1,77 @@
|
||||
---
|
||||
name: autonomous-web-research-agent
|
||||
version: 1.0.0
|
||||
description: Autonomously research a given query on the web using multiple search
|
||||
tools and generate a structured report with summary, detailed sections, source tracking,
|
||||
and bias analysis.
|
||||
inputs:
|
||||
- 'query (string): the research question or topic to investigate'
|
||||
- 'tools (list, optional): selected web search/tools to use (e.g., Tavily, Google,
|
||||
NewsAPI, DuckDuckGo)'
|
||||
- 'api_keys (dict, optional): credentials for LLM and external search APIs'
|
||||
- 'model_config (dict, optional): LLM provider and parameters'
|
||||
steps:
|
||||
- 1. Accept user query and optional tool selections.
|
||||
- '2. Initialize agent framework (e.g., LangGraph) with integrated tools: web search
|
||||
(Tavily, Google, DuckDuckGo), news API, web scraping.'
|
||||
- 3. Decompose query into sub-questions if needed and iteratively call tools to gather
|
||||
relevant information.
|
||||
- 4. Extract and deduplicate content from retrieved sources, tracking source metadata
|
||||
(URL, tool used).
|
||||
- 5. Use a large language model to synthesize findings into an executive summary and
|
||||
detailed sections.
|
||||
- 6. Analyze potential biases or limitations of gathered sources.
|
||||
- 7. Compile a structured report object (ResearchReport) containing query, summary,
|
||||
sections, sources, biases.
|
||||
- 8. Optionally present report via a UI (e.g., Streamlit) or return as JSON.
|
||||
outputs:
|
||||
- 'ResearchReport (JSON/dict) with fields: query (string), summary (string), sections
|
||||
(list of {heading, content}), sources (list of {url, tool_used, title}), potential_biases
|
||||
(string)'
|
||||
- Optional UI rendering of report with source badges and expandable sections
|
||||
tags: []
|
||||
metadata:
|
||||
source_repo: https://github.com/DennisDRX/Faraday-Web-Researcher-Agent.git
|
||||
extracted_at: ''
|
||||
confidence: 0.85
|
||||
---
|
||||
|
||||
# autonomous-web-research-agent
|
||||
|
||||
Autonomously research a given query on the web using multiple search tools and generate a structured report with summary, detailed sections, source tracking, and bias analysis.
|
||||
|
||||
## Steps
|
||||
|
||||
1. 1. Accept user query and optional tool selections.
|
||||
2. 2. Initialize agent framework (e.g., LangGraph) with integrated tools: web search (Tavily, Google, DuckDuckGo), news API, web scraping.
|
||||
3. 3. Decompose query into sub-questions if needed and iteratively call tools to gather relevant information.
|
||||
4. 4. Extract and deduplicate content from retrieved sources, tracking source metadata (URL, tool used).
|
||||
5. 5. Use a large language model to synthesize findings into an executive summary and detailed sections.
|
||||
6. 6. Analyze potential biases or limitations of gathered sources.
|
||||
7. 7. Compile a structured report object (ResearchReport) containing query, summary, sections, sources, biases.
|
||||
8. 8. Optionally present report via a UI (e.g., Streamlit) or return as JSON.
|
||||
|
||||
## Inputs
|
||||
|
||||
- query (string): the research question or topic to investigate
|
||||
- tools (list, optional): selected web search/tools to use (e.g., Tavily, Google, NewsAPI, DuckDuckGo)
|
||||
- api_keys (dict, optional): credentials for LLM and external search APIs
|
||||
- model_config (dict, optional): LLM provider and parameters
|
||||
|
||||
## Outputs
|
||||
|
||||
- ResearchReport (JSON/dict) with fields: query (string), summary (string), sections (list of {heading, content}), sources (list of {url, tool_used, title}), potential_biases (string)
|
||||
- Optional UI rendering of report with source badges and expandable sections
|
||||
|
||||
## Failure Modes
|
||||
|
||||
- Missing or invalid API keys causing tool authentication failures
|
||||
- Rate limits or network errors from search APIs
|
||||
- Insufficient or low-quality search results leading to incomplete report
|
||||
- LLM hallucination or mis-summarization despite source tracking
|
||||
- Parsing errors in HTML/scraped content
|
||||
|
||||
## Source
|
||||
|
||||
Extracted from: [https://github.com/DennisDRX/Faraday-Web-Researcher-Agent.git](https://github.com/DennisDRX/Faraday-Web-Researcher-Agent.git)
|
||||
Confidence: 0.85
|
||||
@@ -0,0 +1,6 @@
|
||||
# Commands: autonomous-web-research-agent
|
||||
|
||||
## Available Commands
|
||||
|
||||
- `/skill autonomous-web-research-agent` — Load this skill
|
||||
- `/run autonomous-web-research-agent` — Execute workflow
|
||||
@@ -0,0 +1,10 @@
|
||||
# Examples: autonomous-web-research-agent
|
||||
|
||||
## Usage Example
|
||||
|
||||
```python
|
||||
# How to use this skill
|
||||
# Inputs: query (string): the research question or topic to investigate, tools (list, optional): selected web search/tools to use (e.g., Tavily, Google, NewsAPI, DuckDuckGo), api_keys (dict, optional): credentials for LLM and external search APIs, model_config (dict, optional): LLM provider and parameters
|
||||
# Process: 1. Accept user query and optional tool selections. → 2. Initialize agent framework (e.g., LangGraph) with integrated tools: web search (Tavily, Google, DuckDuckGo), news API, web scraping. → 3. Decompose query into sub-questions if needed and iteratively call tools to gather relevant information.
|
||||
# Outputs: ResearchReport (JSON/dict) with fields: query (string), summary (string), sections (list of {heading, content}), sources (list of {url, tool_used, title}), potential_biases (string), Optional UI rendering of report with source badges and expandable sections
|
||||
```
|
||||
@@ -0,0 +1,36 @@
|
||||
{
|
||||
"name": "autonomous-web-research-agent",
|
||||
"version": "1.0.0",
|
||||
"goal": "Autonomously research a given query on the web using multiple search tools and generate a structured report with summary, detailed sections, source tracking, and bias analysis.",
|
||||
"inputs": [
|
||||
"query (string): the research question or topic to investigate",
|
||||
"tools (list, optional): selected web search/tools to use (e.g., Tavily, Google, NewsAPI, DuckDuckGo)",
|
||||
"api_keys (dict, optional): credentials for LLM and external search APIs",
|
||||
"model_config (dict, optional): LLM provider and parameters"
|
||||
],
|
||||
"steps": [
|
||||
"1. Accept user query and optional tool selections.",
|
||||
"2. Initialize agent framework (e.g., LangGraph) with integrated tools: web search (Tavily, Google, DuckDuckGo), news API, web scraping.",
|
||||
"3. Decompose query into sub-questions if needed and iteratively call tools to gather relevant information.",
|
||||
"4. Extract and deduplicate content from retrieved sources, tracking source metadata (URL, tool used).",
|
||||
"5. Use a large language model to synthesize findings into an executive summary and detailed sections.",
|
||||
"6. Analyze potential biases or limitations of gathered sources.",
|
||||
"7. Compile a structured report object (ResearchReport) containing query, summary, sections, sources, biases.",
|
||||
"8. Optionally present report via a UI (e.g., Streamlit) or return as JSON."
|
||||
],
|
||||
"outputs": [
|
||||
"ResearchReport (JSON/dict) with fields: query (string), summary (string), sections (list of {heading, content}), sources (list of {url, tool_used, title}), potential_biases (string)",
|
||||
"Optional UI rendering of report with source badges and expandable sections"
|
||||
],
|
||||
"failure_modes": [
|
||||
"Missing or invalid API keys causing tool authentication failures",
|
||||
"Rate limits or network errors from search APIs",
|
||||
"Insufficient or low-quality search results leading to incomplete report",
|
||||
"LLM hallucination or mis-summarization despite source tracking",
|
||||
"Parsing errors in HTML/scraped content"
|
||||
],
|
||||
"confidence": 0.85,
|
||||
"explanation": "The repository implements a generic autonomous web research agent that can be reused for any topical query. The workflow of querying, multi-tool retrieval, synthesis, and structured reporting is not domain-specific and can be extracted as a reusable skill.",
|
||||
"source_repo": "https://github.com/DennisDRX/Faraday-Web-Researcher-Agent.git",
|
||||
"score": 1.0
|
||||
}
|
||||
+1
-1
@@ -1,4 +1,4 @@
|
||||
# Tests: local-document-research-with-traceable-citations
|
||||
# Tests: autonomous-web-research-agent
|
||||
|
||||
## Test Checklist
|
||||
|
||||
@@ -1,82 +0,0 @@
|
||||
---
|
||||
name: local-document-research-with-traceable-citations
|
||||
version: 1.0.0
|
||||
description: Enable users to import local documents, asynchronously process them into
|
||||
an indexed knowledge base, and obtain AI-generated answers that cite specific page
|
||||
locations and OCR evidence.
|
||||
inputs:
|
||||
- Local document files (PDF, DOCX, PPTX, XLSX, images, TXT)
|
||||
- Configured LLM API endpoint and keys (via .env or settings)
|
||||
- Optional web search service config if enabled
|
||||
- Local OCR model cache (downloaded on first use)
|
||||
steps:
|
||||
- Import documents into a project; files are queued for asynchronous processing.
|
||||
- Convert non-PDF formats (DOCX, PPTX, XLSX) to PDF using native Office or LibreOffice
|
||||
fallback.
|
||||
- Run OCR (PaddleOCR) on PDF pages/images to extract text, page numbers, polygons,
|
||||
and confidence scores.
|
||||
- Chunk text and generate embeddings; publish to LanceDB hybrid index (dense vector
|
||||
+ full-text) only when fully processed.
|
||||
- User starts a research session or branch; Leader agent analyzes query.
|
||||
- Hybrid RAG retrieves candidate chunks from ready documents; dynamic material scope
|
||||
ensures no half-indexed docs.
|
||||
- Leader delegates tasks to sub-agents (researcher, reviewer, writer) via constrained
|
||||
task capability; each delegation logs start, completion, duration, and evidence.
|
||||
- Generate answer that references only actually used evidence; citations include document
|
||||
ID, page, coordinates.
|
||||
- User clicks citation to open original document and view highlighted OCR location.
|
||||
outputs:
|
||||
- Project with indexed document library (SQLite metadata + LanceDB vectors)
|
||||
- AI answers with verifiable citations to source pages
|
||||
- Evidence preview with page image and OCR highlight polygons
|
||||
- Persistent session history, branches, and long-term memory
|
||||
tags: []
|
||||
metadata:
|
||||
source_repo: https://github.com/0verL1nk/PaperSage.git
|
||||
extracted_at: ''
|
||||
confidence: 0.85
|
||||
---
|
||||
|
||||
# local-document-research-with-traceable-citations
|
||||
|
||||
Enable users to import local documents, asynchronously process them into an indexed knowledge base, and obtain AI-generated answers that cite specific page locations and OCR evidence.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Import documents into a project; files are queued for asynchronous processing.
|
||||
2. Convert non-PDF formats (DOCX, PPTX, XLSX) to PDF using native Office or LibreOffice fallback.
|
||||
3. Run OCR (PaddleOCR) on PDF pages/images to extract text, page numbers, polygons, and confidence scores.
|
||||
4. Chunk text and generate embeddings; publish to LanceDB hybrid index (dense vector + full-text) only when fully processed.
|
||||
5. User starts a research session or branch; Leader agent analyzes query.
|
||||
6. Hybrid RAG retrieves candidate chunks from ready documents; dynamic material scope ensures no half-indexed docs.
|
||||
7. Leader delegates tasks to sub-agents (researcher, reviewer, writer) via constrained task capability; each delegation logs start, completion, duration, and evidence.
|
||||
8. Generate answer that references only actually used evidence; citations include document ID, page, coordinates.
|
||||
9. User clicks citation to open original document and view highlighted OCR location.
|
||||
|
||||
## Inputs
|
||||
|
||||
- Local document files (PDF, DOCX, PPTX, XLSX, images, TXT)
|
||||
- Configured LLM API endpoint and keys (via .env or settings)
|
||||
- Optional web search service config if enabled
|
||||
- Local OCR model cache (downloaded on first use)
|
||||
|
||||
## Outputs
|
||||
|
||||
- Project with indexed document library (SQLite metadata + LanceDB vectors)
|
||||
- AI answers with verifiable citations to source pages
|
||||
- Evidence preview with page image and OCR highlight polygons
|
||||
- Persistent session history, branches, and long-term memory
|
||||
|
||||
## Failure Modes
|
||||
|
||||
- OCR quality low for scanned images leading to poor extraction
|
||||
- Missing Office/LibreOffice causes conversion failure for Office docs
|
||||
- Interrupted indexing leaves documents unpublished and excluded from retrieval
|
||||
- LLM API outage or misconfiguration yields no answer
|
||||
- Citation coordinates mismatch due to chunk drift
|
||||
- Sub-agent recursion if constraints not enforced
|
||||
|
||||
## Source
|
||||
|
||||
Extracted from: [https://github.com/0verL1nk/PaperSage.git](https://github.com/0verL1nk/PaperSage.git)
|
||||
Confidence: 0.85
|
||||
@@ -1,6 +0,0 @@
|
||||
# Commands: local-document-research-with-traceable-citations
|
||||
|
||||
## Available Commands
|
||||
|
||||
- `/skill local-document-research-with-traceable-citations` — Load this skill
|
||||
- `/run local-document-research-with-traceable-citations` — Execute workflow
|
||||
@@ -1,10 +0,0 @@
|
||||
# Examples: local-document-research-with-traceable-citations
|
||||
|
||||
## Usage Example
|
||||
|
||||
```python
|
||||
# How to use this skill
|
||||
# Inputs: Local document files (PDF, DOCX, PPTX, XLSX, images, TXT), Configured LLM API endpoint and keys (via .env or settings), Optional web search service config if enabled, Local OCR model cache (downloaded on first use)
|
||||
# Process: Import documents into a project; files are queued for asynchronous processing. → Convert non-PDF formats (DOCX, PPTX, XLSX) to PDF using native Office or LibreOffice fallback. → Run OCR (PaddleOCR) on PDF pages/images to extract text, page numbers, polygons, and confidence scores.
|
||||
# Outputs: Project with indexed document library (SQLite metadata + LanceDB vectors), AI answers with verifiable citations to source pages, Evidence preview with page image and OCR highlight polygons, Persistent session history, branches, and long-term memory
|
||||
```
|
||||
@@ -1,40 +0,0 @@
|
||||
{
|
||||
"name": "local-document-research-with-traceable-citations",
|
||||
"version": "1.0.0",
|
||||
"goal": "Enable users to import local documents, asynchronously process them into an indexed knowledge base, and obtain AI-generated answers that cite specific page locations and OCR evidence.",
|
||||
"inputs": [
|
||||
"Local document files (PDF, DOCX, PPTX, XLSX, images, TXT)",
|
||||
"Configured LLM API endpoint and keys (via .env or settings)",
|
||||
"Optional web search service config if enabled",
|
||||
"Local OCR model cache (downloaded on first use)"
|
||||
],
|
||||
"steps": [
|
||||
"Import documents into a project; files are queued for asynchronous processing.",
|
||||
"Convert non-PDF formats (DOCX, PPTX, XLSX) to PDF using native Office or LibreOffice fallback.",
|
||||
"Run OCR (PaddleOCR) on PDF pages/images to extract text, page numbers, polygons, and confidence scores.",
|
||||
"Chunk text and generate embeddings; publish to LanceDB hybrid index (dense vector + full-text) only when fully processed.",
|
||||
"User starts a research session or branch; Leader agent analyzes query.",
|
||||
"Hybrid RAG retrieves candidate chunks from ready documents; dynamic material scope ensures no half-indexed docs.",
|
||||
"Leader delegates tasks to sub-agents (researcher, reviewer, writer) via constrained task capability; each delegation logs start, completion, duration, and evidence.",
|
||||
"Generate answer that references only actually used evidence; citations include document ID, page, coordinates.",
|
||||
"User clicks citation to open original document and view highlighted OCR location."
|
||||
],
|
||||
"outputs": [
|
||||
"Project with indexed document library (SQLite metadata + LanceDB vectors)",
|
||||
"AI answers with verifiable citations to source pages",
|
||||
"Evidence preview with page image and OCR highlight polygons",
|
||||
"Persistent session history, branches, and long-term memory"
|
||||
],
|
||||
"failure_modes": [
|
||||
"OCR quality low for scanned images leading to poor extraction",
|
||||
"Missing Office/LibreOffice causes conversion failure for Office docs",
|
||||
"Interrupted indexing leaves documents unpublished and excluded from retrieval",
|
||||
"LLM API outage or misconfiguration yields no answer",
|
||||
"Citation coordinates mismatch due to chunk drift",
|
||||
"Sub-agent recursion if constraints not enforced"
|
||||
],
|
||||
"confidence": 0.85,
|
||||
"explanation": "The README outlines a clear pipeline from document import to cited answer with evidence location, which is a reusable pattern for local-first RAG applications requiring traceability.",
|
||||
"source_repo": "https://github.com/0verL1nk/PaperSage.git",
|
||||
"score": 1.0
|
||||
}
|
||||
Reference in New Issue
Block a user