Compare commits

..

1 Commits

Author SHA1 Message Date
Hermes Pipeline 28b0c483b8 Add Skill: autonomous-web-research-agent
Extracted from: https://github.com/DennisDRX/Faraday-Web-Researcher-Agent.git
Score: 1.0
2026-08-10 17:05:01 +00:00
9 changed files with 130 additions and 131 deletions
@@ -0,0 +1,77 @@
---
name: autonomous-web-research-agent
version: 1.0.0
description: Autonomously research a given query on the web using multiple search
tools and generate a structured report with summary, detailed sections, source tracking,
and bias analysis.
inputs:
- 'query (string): the research question or topic to investigate'
- 'tools (list, optional): selected web search/tools to use (e.g., Tavily, Google,
NewsAPI, DuckDuckGo)'
- 'api_keys (dict, optional): credentials for LLM and external search APIs'
- 'model_config (dict, optional): LLM provider and parameters'
steps:
- 1. Accept user query and optional tool selections.
- '2. Initialize agent framework (e.g., LangGraph) with integrated tools: web search
(Tavily, Google, DuckDuckGo), news API, web scraping.'
- 3. Decompose query into sub-questions if needed and iteratively call tools to gather
relevant information.
- 4. Extract and deduplicate content from retrieved sources, tracking source metadata
(URL, tool used).
- 5. Use a large language model to synthesize findings into an executive summary and
detailed sections.
- 6. Analyze potential biases or limitations of gathered sources.
- 7. Compile a structured report object (ResearchReport) containing query, summary,
sections, sources, biases.
- 8. Optionally present report via a UI (e.g., Streamlit) or return as JSON.
outputs:
- 'ResearchReport (JSON/dict) with fields: query (string), summary (string), sections
(list of {heading, content}), sources (list of {url, tool_used, title}), potential_biases
(string)'
- Optional UI rendering of report with source badges and expandable sections
tags: []
metadata:
source_repo: https://github.com/DennisDRX/Faraday-Web-Researcher-Agent.git
extracted_at: ''
confidence: 0.85
---
# autonomous-web-research-agent
Autonomously research a given query on the web using multiple search tools and generate a structured report with summary, detailed sections, source tracking, and bias analysis.
## Steps
1. 1. Accept user query and optional tool selections.
2. 2. Initialize agent framework (e.g., LangGraph) with integrated tools: web search (Tavily, Google, DuckDuckGo), news API, web scraping.
3. 3. Decompose query into sub-questions if needed and iteratively call tools to gather relevant information.
4. 4. Extract and deduplicate content from retrieved sources, tracking source metadata (URL, tool used).
5. 5. Use a large language model to synthesize findings into an executive summary and detailed sections.
6. 6. Analyze potential biases or limitations of gathered sources.
7. 7. Compile a structured report object (ResearchReport) containing query, summary, sections, sources, biases.
8. 8. Optionally present report via a UI (e.g., Streamlit) or return as JSON.
## Inputs
- query (string): the research question or topic to investigate
- tools (list, optional): selected web search/tools to use (e.g., Tavily, Google, NewsAPI, DuckDuckGo)
- api_keys (dict, optional): credentials for LLM and external search APIs
- model_config (dict, optional): LLM provider and parameters
## Outputs
- ResearchReport (JSON/dict) with fields: query (string), summary (string), sections (list of {heading, content}), sources (list of {url, tool_used, title}), potential_biases (string)
- Optional UI rendering of report with source badges and expandable sections
## Failure Modes
- Missing or invalid API keys causing tool authentication failures
- Rate limits or network errors from search APIs
- Insufficient or low-quality search results leading to incomplete report
- LLM hallucination or mis-summarization despite source tracking
- Parsing errors in HTML/scraped content
## Source
Extracted from: [https://github.com/DennisDRX/Faraday-Web-Researcher-Agent.git](https://github.com/DennisDRX/Faraday-Web-Researcher-Agent.git)
Confidence: 0.85
@@ -0,0 +1,6 @@
# Commands: autonomous-web-research-agent
## Available Commands
- `/skill autonomous-web-research-agent` — Load this skill
- `/run autonomous-web-research-agent` — Execute workflow
@@ -0,0 +1,10 @@
# Examples: autonomous-web-research-agent
## Usage Example
```python
# How to use this skill
# Inputs: query (string): the research question or topic to investigate, tools (list, optional): selected web search/tools to use (e.g., Tavily, Google, NewsAPI, DuckDuckGo), api_keys (dict, optional): credentials for LLM and external search APIs, model_config (dict, optional): LLM provider and parameters
# Process: 1. Accept user query and optional tool selections. → 2. Initialize agent framework (e.g., LangGraph) with integrated tools: web search (Tavily, Google, DuckDuckGo), news API, web scraping. → 3. Decompose query into sub-questions if needed and iteratively call tools to gather relevant information.
# Outputs: ResearchReport (JSON/dict) with fields: query (string), summary (string), sections (list of {heading, content}), sources (list of {url, tool_used, title}), potential_biases (string), Optional UI rendering of report with source badges and expandable sections
```
@@ -0,0 +1,36 @@
{
"name": "autonomous-web-research-agent",
"version": "1.0.0",
"goal": "Autonomously research a given query on the web using multiple search tools and generate a structured report with summary, detailed sections, source tracking, and bias analysis.",
"inputs": [
"query (string): the research question or topic to investigate",
"tools (list, optional): selected web search/tools to use (e.g., Tavily, Google, NewsAPI, DuckDuckGo)",
"api_keys (dict, optional): credentials for LLM and external search APIs",
"model_config (dict, optional): LLM provider and parameters"
],
"steps": [
"1. Accept user query and optional tool selections.",
"2. Initialize agent framework (e.g., LangGraph) with integrated tools: web search (Tavily, Google, DuckDuckGo), news API, web scraping.",
"3. Decompose query into sub-questions if needed and iteratively call tools to gather relevant information.",
"4. Extract and deduplicate content from retrieved sources, tracking source metadata (URL, tool used).",
"5. Use a large language model to synthesize findings into an executive summary and detailed sections.",
"6. Analyze potential biases or limitations of gathered sources.",
"7. Compile a structured report object (ResearchReport) containing query, summary, sections, sources, biases.",
"8. Optionally present report via a UI (e.g., Streamlit) or return as JSON."
],
"outputs": [
"ResearchReport (JSON/dict) with fields: query (string), summary (string), sections (list of {heading, content}), sources (list of {url, tool_used, title}), potential_biases (string)",
"Optional UI rendering of report with source badges and expandable sections"
],
"failure_modes": [
"Missing or invalid API keys causing tool authentication failures",
"Rate limits or network errors from search APIs",
"Insufficient or low-quality search results leading to incomplete report",
"LLM hallucination or mis-summarization despite source tracking",
"Parsing errors in HTML/scraped content"
],
"confidence": 0.85,
"explanation": "The repository implements a generic autonomous web research agent that can be reused for any topical query. The workflow of querying, multi-tool retrieval, synthesis, and structured reporting is not domain-specific and can be extracted as a reusable skill.",
"source_repo": "https://github.com/DennisDRX/Faraday-Web-Researcher-Agent.git",
"score": 1.0
}
@@ -1,4 +1,4 @@
# Tests: literature-review-with-traceable-ai-evidence
# Tests: autonomous-web-research-agent
## Test Checklist
@@ -1,77 +0,0 @@
---
name: literature-review-with-traceable-ai-evidence
version: 1.0.0
description: Enable researchers to ingest documents, asynchronously index them, and
interact with multi-agent AI to answer questions with verifiable citations to original
text.
inputs:
- Documents in PDF, Office, image, or text formats
- Research questions or topics of interest
- 'Optional: user model configuration via .env or settings'
steps:
- Upload documents to a project (via desktop app or web UI)
- System asynchronously converts Office docs to PDF if needed, runs OCR to extract
text with coordinates, chunks and embeds into vector store
- User starts a main research session or creates exploration branches without waiting
for indexing to finish
- Leader agent receives query and delegates subtasks to researcher, reviewer, writer
subagents
- Subagents perform hybrid retrieval and rerank to find relevant chunks with source
coordinates
- Agents synthesize answers and return evidence with clickable citations that highlight
original pages
- User verifies conclusions by navigating to cited source locations and can save notes
to research memory or mind map
outputs:
- AI-generated answers with traceable evidence (coordinates, page highlights)
- Research session history with branches
- Indexed document library for future queries
- Mind maps or structured notes
- Persistent run events for resuming sessions
tags: []
metadata:
source_repo: https://github.com/0verL1nk/PaperSage.git
extracted_at: ''
confidence: 0.85
---
# literature-review-with-traceable-ai-evidence
Enable researchers to ingest documents, asynchronously index them, and interact with multi-agent AI to answer questions with verifiable citations to original text.
## Steps
1. Upload documents to a project (via desktop app or web UI)
2. System asynchronously converts Office docs to PDF if needed, runs OCR to extract text with coordinates, chunks and embeds into vector store
3. User starts a main research session or creates exploration branches without waiting for indexing to finish
4. Leader agent receives query and delegates subtasks to researcher, reviewer, writer subagents
5. Subagents perform hybrid retrieval and rerank to find relevant chunks with source coordinates
6. Agents synthesize answers and return evidence with clickable citations that highlight original pages
7. User verifies conclusions by navigating to cited source locations and can save notes to research memory or mind map
## Inputs
- Documents in PDF, Office, image, or text formats
- Research questions or topics of interest
- Optional: user model configuration via .env or settings
## Outputs
- AI-generated answers with traceable evidence (coordinates, page highlights)
- Research session history with branches
- Indexed document library for future queries
- Mind maps or structured notes
- Persistent run events for resuming sessions
## Failure Modes
- Missing local Office/LibreOffice converter causes document conversion failure
- First-time model download may be slow or require network
- OCR may have low confidence on poor quality scans
- Retrieval might miss context if chunking splits semantics
- Multi-agent coordination could produce conflicting intermediate results
## Source
Extracted from: [https://github.com/0verL1nk/PaperSage.git](https://github.com/0verL1nk/PaperSage.git)
Confidence: 0.85
@@ -1,6 +0,0 @@
# Commands: literature-review-with-traceable-ai-evidence
## Available Commands
- `/skill literature-review-with-traceable-ai-evidence` — Load this skill
- `/run literature-review-with-traceable-ai-evidence` — Execute workflow
@@ -1,10 +0,0 @@
# Examples: literature-review-with-traceable-ai-evidence
## Usage Example
```python
# How to use this skill
# Inputs: Documents in PDF, Office, image, or text formats, Research questions or topics of interest, Optional: user model configuration via .env or settings
# Process: Upload documents to a project (via desktop app or web UI) → System asynchronously converts Office docs to PDF if needed, runs OCR to extract text with coordinates, chunks and embeds into vector store → User starts a main research session or creates exploration branches without waiting for indexing to finish
# Outputs: AI-generated answers with traceable evidence (coordinates, page highlights), Research session history with branches, Indexed document library for future queries, Mind maps or structured notes, Persistent run events for resuming sessions
```
@@ -1,37 +0,0 @@
{
"name": "literature-review-with-traceable-ai-evidence",
"version": "1.0.0",
"goal": "Enable researchers to ingest documents, asynchronously index them, and interact with multi-agent AI to answer questions with verifiable citations to original text.",
"inputs": [
"Documents in PDF, Office, image, or text formats",
"Research questions or topics of interest",
"Optional: user model configuration via .env or settings"
],
"steps": [
"Upload documents to a project (via desktop app or web UI)",
"System asynchronously converts Office docs to PDF if needed, runs OCR to extract text with coordinates, chunks and embeds into vector store",
"User starts a main research session or creates exploration branches without waiting for indexing to finish",
"Leader agent receives query and delegates subtasks to researcher, reviewer, writer subagents",
"Subagents perform hybrid retrieval and rerank to find relevant chunks with source coordinates",
"Agents synthesize answers and return evidence with clickable citations that highlight original pages",
"User verifies conclusions by navigating to cited source locations and can save notes to research memory or mind map"
],
"outputs": [
"AI-generated answers with traceable evidence (coordinates, page highlights)",
"Research session history with branches",
"Indexed document library for future queries",
"Mind maps or structured notes",
"Persistent run events for resuming sessions"
],
"failure_modes": [
"Missing local Office/LibreOffice converter causes document conversion failure",
"First-time model download may be slow or require network",
"OCR may have low confidence on poor quality scans",
"Retrieval might miss context if chunking splits semantics",
"Multi-agent coordination could produce conflicting intermediate results"
],
"confidence": 0.85,
"explanation": "The README describes PaperSage's core workflow: asynchronous document ingestion with OCR/indexing, followed by multi-agent question answering with cited evidence. This process is not tied to the specific codebase and can be reused as a general literature review methodology for any document-centric research using RAG and agent collaboration.",
"source_repo": "https://github.com/0verL1nk/PaperSage.git",
"score": 1.0
}