Compare commits

..

4 Commits

Author SHA1 Message Date
Hermes Pipeline 1ba56fd7e3 Add Skill: three-tier-evaluation-pipeline
Extracted from: https://github.com/itszhaoziyan-n/AgentKit.git
Score: 1.0
2026-08-05 17:05:03 +00:00
Epictetus d81ddeda88 Scout window: 6 months (Feb-Aug 2026)
- pushed_after: 2026-02-01 (was 2024-06-01)
- Keeps scope tight on recent, relevant repos
2026-08-05 16:44:22 +00:00
Epictetus a14f09bec2 Add 2 new skills + dedup fix
New skills:
- langgraph-workflow-creation (from LangGraphProjects)
- agent-supervisor (from multi_agent_workflow_demo_in_langgraph)

Fixes:
- Publisher dedup: skip existing skills
- Older repos: pushed_after 2024-06-01 (was 2026-05-01)
- Lower stars: 10 (was 15)
2026-08-05 15:55:36 +00:00
Epictetus 7f496feb90 Publisher dedup check + run.py skip display
- Skip skills already in skills/ directory (no duplicate PRs)
- Run.py shows SKIP status with reason
- Fixed: was re-publishing same 5 skills every run
2026-08-05 15:49:57 +00:00
21 changed files with 530 additions and 61 deletions
+2 -2
View File
@@ -34,8 +34,8 @@ scout:
- 'rag agent workflow'
- 'tool calling workflow'
filters:
stars_min: 15
pushed_after: 2026-05-01
stars_min: 10
pushed_after: 2026-02-01
language: Python
archived: false
size_max_kb: 10000
+12
View File
@@ -52,6 +52,18 @@ def publish_skill(review_result, config):
subprocess.run(["git", "config", "user.email", "hermes@agent.local"], cwd=repo_dir)
subprocess.run(["git", "config", "user.name", "Hermes Pipeline"], cwd=repo_dir)
# Check for duplicates in skills/ directory
skills_dir = os.path.join(repo_dir, "skills")
existing_skills = []
if os.path.isdir(skills_dir):
existing_skills = [d for d in os.listdir(skills_dir) if os.path.isdir(os.path.join(skills_dir, d))]
if skill_name in existing_skills:
return {
"status": "SKIP",
"reason": f"Skill '{skill_name}' already exists in skills/ directory",
}
# Create skill directory
skill_dir = os.path.join(repo_dir, "skills", skill_name)
os.makedirs(skill_dir, exist_ok=True)
+2
View File
@@ -137,6 +137,8 @@ def main():
if publish_output.get("status") == "PUBLISHED":
print(f" ✓ Published! PR: {publish_output.get('pr_url', '')}")
results["published"] += 1
elif publish_output.get("status") == "SKIP":
print(f" ⏸ Skipped: {publish_output.get('reason', '')}")
else:
print(f" ! {publish_output.get('status', '?')}: {publish_output.get('message', publish_output.get('error', ''))[:100]}")
+84
View File
@@ -0,0 +1,84 @@
---
name: agent-supervisor
version: 1.0.0
description: Demonstrate a supervisor-worker architecture for intelligent task delegation
and real-time decision-making.
inputs:
- name: OPENAI_API_KEY
description: OpenAI API key for language models.
- name: TAVILY_API_KEY
description: Tavily API key for search functionality.
steps:
- step: 1
action: Load environment variables.
details: Set the OPENAI_API_KEY and TAVILY_API_KEY environment variables.
- step: 2
action: Configure LangChain tools.
details: Initialize TavilySearchResults and PythonREPLTool.
- step: 3
action: Define agent nodes.
details: Create functions for the Researcher and Coder agents that process state
through their respective tasks.
- step: 4
action: Set up supervisor agent.
details: Create a supervisor agent function that decides which worker should act
next based on user input.
- step: 5
action: Build state graph.
details: Construct the state graph with nodes for each agent and edges connecting
them to the supervisor node.
- step: 6
action: Add conditional edges.
details: Define conditions for transitioning between agents based on their responses.
- step: 7
action: Compile graph.
details: Compile the state graph into a runnable workflow.
- step: 8
action: Run example queries.
details: Stream through the workflow with example inputs to demonstrate its functionality.
outputs:
- name: 'Example 1: Code Hello World'
description: A demonstration of coding a simple hello world program.
- name: 'Example 2: Research Report'
description: A demonstration of researching and writing a brief report on pikas.
tags: []
metadata:
source_repo: https://github.com/extrawest/multi_agent_workflow_demo_in_langgraph.git
extracted_at: ''
confidence: 0.9
---
# agent-supervisor
Demonstrate a supervisor-worker architecture for intelligent task delegation and real-time decision-making.
## Steps
1. {'step': 1, 'action': 'Load environment variables.', 'details': 'Set the OPENAI_API_KEY and TAVILY_API_KEY environment variables.'}
2. {'step': 2, 'action': 'Configure LangChain tools.', 'details': 'Initialize TavilySearchResults and PythonREPLTool.'}
3. {'step': 3, 'action': 'Define agent nodes.', 'details': 'Create functions for the Researcher and Coder agents that process state through their respective tasks.'}
4. {'step': 4, 'action': 'Set up supervisor agent.', 'details': 'Create a supervisor agent function that decides which worker should act next based on user input.'}
5. {'step': 5, 'action': 'Build state graph.', 'details': 'Construct the state graph with nodes for each agent and edges connecting them to the supervisor node.'}
6. {'step': 6, 'action': 'Add conditional edges.', 'details': 'Define conditions for transitioning between agents based on their responses.'}
7. {'step': 7, 'action': 'Compile graph.', 'details': 'Compile the state graph into a runnable workflow.'}
8. {'step': 8, 'action': 'Run example queries.', 'details': 'Stream through the workflow with example inputs to demonstrate its functionality.'}
## Inputs
- {'name': 'OPENAI_API_KEY', 'description': 'OpenAI API key for language models.'}
- {'name': 'TAVILY_API_KEY', 'description': 'Tavily API key for search functionality.'}
## Outputs
- {'name': 'Example 1: Code Hello World', 'description': 'A demonstration of coding a simple hello world program.'}
- {'name': 'Example 2: Research Report', 'description': 'A demonstration of researching and writing a brief report on pikas.'}
## Failure Modes
- {'mode': 'Invalid API keys', 'description': 'The workflow may fail if the provided API keys are invalid or expired.'}
- {'mode': 'Insufficient permissions', 'description': 'The workflow may fail if the user does not have sufficient permissions to use the Tavily search functionality.'}
## Source
Extracted from: [https://github.com/extrawest/multi_agent_workflow_demo_in_langgraph.git](https://github.com/extrawest/multi_agent_workflow_demo_in_langgraph.git)
Confidence: 0.9
+6
View File
@@ -0,0 +1,6 @@
# Commands: agent-supervisor
## Available Commands
- `/skill agent-supervisor` — Load this skill
- `/run agent-supervisor` — Execute workflow
+10
View File
@@ -0,0 +1,10 @@
# Examples: agent-supervisor
## Usage Example
```python
# How to use this skill
# Inputs: {'name': 'OPENAI_API_KEY', 'description': 'OpenAI API key for language models.'}, {'name': 'TAVILY_API_KEY', 'description': 'Tavily API key for search functionality.'}
# Process: {'step': 1, 'action': 'Load environment variables.', 'details': 'Set the OPENAI_API_KEY and TAVILY_API_KEY environment variables.'} → {'step': 2, 'action': 'Configure LangChain tools.', 'details': 'Initialize TavilySearchResults and PythonREPLTool.'} → {'step': 3, 'action': 'Define agent nodes.', 'details': 'Create functions for the Researcher and Coder agents that process state through their respective tasks.'}
# Outputs: {'name': 'Example 1: Code Hello World', 'description': 'A demonstration of coding a simple hello world program.'}, {'name': 'Example 2: Research Report', 'description': 'A demonstration of researching and writing a brief report on pikas.'}
```
+81
View File
@@ -0,0 +1,81 @@
{
"name": "agent-supervisor",
"version": "1.0.0",
"goal": "Demonstrate a supervisor-worker architecture for intelligent task delegation and real-time decision-making.",
"inputs": [
{
"name": "OPENAI_API_KEY",
"description": "OpenAI API key for language models."
},
{
"name": "TAVILY_API_KEY",
"description": "Tavily API key for search functionality."
}
],
"steps": [
{
"step": 1,
"action": "Load environment variables.",
"details": "Set the OPENAI_API_KEY and TAVILY_API_KEY environment variables."
},
{
"step": 2,
"action": "Configure LangChain tools.",
"details": "Initialize TavilySearchResults and PythonREPLTool."
},
{
"step": 3,
"action": "Define agent nodes.",
"details": "Create functions for the Researcher and Coder agents that process state through their respective tasks."
},
{
"step": 4,
"action": "Set up supervisor agent.",
"details": "Create a supervisor agent function that decides which worker should act next based on user input."
},
{
"step": 5,
"action": "Build state graph.",
"details": "Construct the state graph with nodes for each agent and edges connecting them to the supervisor node."
},
{
"step": 6,
"action": "Add conditional edges.",
"details": "Define conditions for transitioning between agents based on their responses."
},
{
"step": 7,
"action": "Compile graph.",
"details": "Compile the state graph into a runnable workflow."
},
{
"step": 8,
"action": "Run example queries.",
"details": "Stream through the workflow with example inputs to demonstrate its functionality."
}
],
"outputs": [
{
"name": "Example 1: Code Hello World",
"description": "A demonstration of coding a simple hello world program."
},
{
"name": "Example 2: Research Report",
"description": "A demonstration of researching and writing a brief report on pikas."
}
],
"failure_modes": [
{
"mode": "Invalid API keys",
"description": "The workflow may fail if the provided API keys are invalid or expired."
},
{
"mode": "Insufficient permissions",
"description": "The workflow may fail if the user does not have sufficient permissions to use the Tavily search functionality."
}
],
"confidence": 0.9,
"explanation": "This workflow demonstrates a hierarchical multi-agent system where a supervisor agent makes routing decisions based on user input, delegating tasks to specialized worker agents (Researcher and Coder). It is designed to be reusable for similar task delegation scenarios.",
"source_repo": "https://github.com/extrawest/multi_agent_workflow_demo_in_langgraph.git",
"score": 1.0
}
+9
View File
@@ -0,0 +1,9 @@
# Tests: agent-supervisor
## Test Checklist
- [ ] Workflow has at least 3 steps
- [ ] All inputs are defined
- [ ] All outputs are defined
- [ ] Failure modes are documented
- [ ] Skill can be loaded without errors
@@ -0,0 +1,83 @@
---
name: langgraph-workflow-creation
version: 1.0.0
description: Create a LangGraph workflow to gather facts using SerperDevTool and process
them with an AI agent.
inputs:
- API Key for SerperDevTool
- Search Query
steps:
- 'Step 1: Import necessary modules from langgraph and langchain libraries'
- 'Step 2: Create a LangGraph agent using `langgraph.create_react_agent` function
with SerperDevTool as the tool node'
- 'Step 3: Define the search query and pass it to the agent for fact gathering'
- 'Step 4: Process the gathered facts within the AI agent'
outputs:
- Processed Facts
tags: []
metadata:
source_repo: https://github.com/jkmaina/LangGraphProjects.git
extracted_at: ''
confidence: 0.95
---
# langgraph-workflow-creation
Create a LangGraph workflow to gather facts using SerperDevTool and process them with an AI agent.
## Setup
**Dependencies:**
```text
pip install langchain serperdev
```
**Setup steps:**
1. Install required libraries: pip install langchain serperdev
1. Add API key to .env file: OPENAPI_API_KEY=your_api_key
## Key Files
- `agent.py - Contains the LangGraph agent creation logic`
- `tool_node.py - Defines the SerperDevTool node`
## Steps
1. Step 1: Import necessary modules from langgraph and langchain libraries
2. Step 2: Create a LangGraph agent using `langgraph.create_react_agent` function with SerperDevTool as the tool node
3. Step 3: Define the search query and pass it to the agent for fact gathering
4. Step 4: Process the gathered facts within the AI agent
## Implementation Details
```python
import langgraph
from serperdev import SerperDevTool
def create_agent(api_key, query):
tool = SerperDevTool(api_key)
agent = langgraph.create_react_agent(tool=tool)
facts = agent.run(query)
return process_facts(facts)
```
## Inputs
- API Key for SerperDevTool
- Search Query
## Outputs
- Processed Facts
## Failure Modes
- API Key not provided
- Invalid Search Query
## Source
Extracted from: [https://github.com/jkmaina/LangGraphProjects.git](https://github.com/jkmaina/LangGraphProjects.git)
Confidence: 0.95
@@ -0,0 +1,6 @@
# Commands: langgraph-workflow-creation
## Available Commands
- `/skill langgraph-workflow-creation` — Load this skill
- `/run langgraph-workflow-creation` — Execute workflow
@@ -0,0 +1,10 @@
# Examples: langgraph-workflow-creation
## Usage Example
```python
# How to use this skill
# Inputs: API Key for SerperDevTool, Search Query
# Process: Step 1: Import necessary modules from langgraph and langchain libraries → Step 2: Create a LangGraph agent using `langgraph.create_react_agent` function with SerperDevTool as the tool node → Step 3: Define the search query and pass it to the agent for fact gathering
# Outputs: Processed Facts
```
@@ -0,0 +1,26 @@
{
"name": "langgraph-workflow-creation",
"version": "1.0.0",
"goal": "Create a LangGraph workflow to gather facts using SerperDevTool and process them with an AI agent.",
"inputs": [
"API Key for SerperDevTool",
"Search Query"
],
"steps": [
"Step 1: Import necessary modules from langgraph and langchain libraries",
"Step 2: Create a LangGraph agent using `langgraph.create_react_agent` function with SerperDevTool as the tool node",
"Step 3: Define the search query and pass it to the agent for fact gathering",
"Step 4: Process the gathered facts within the AI agent"
],
"outputs": [
"Processed Facts"
],
"failure_modes": [
"API Key not provided",
"Invalid Search Query"
],
"confidence": 0.95,
"explanation": "This workflow is specific to fact gathering and can be adapted for different search queries or tools.",
"source_repo": "https://github.com/jkmaina/LangGraphProjects.git",
"score": 1.0
}
@@ -0,0 +1,9 @@
# Tests: langgraph-workflow-creation
## Test Checklist
- [ ] Workflow has at least 3 steps
- [ ] All inputs are defined
- [ ] All outputs are defined
- [ ] Failure modes are documented
- [ ] Skill can be loaded without errors
+30 -45
View File
@@ -1,27 +1,26 @@
---
name: research-pipeline
version: 1.0.0
description: Fetch a Wikipedia page, summarise it using an AI agent, and write the
summary to a file.
description: Fetch a Wikipedia page, summarise its content using an AI agent, and
write the summary to a file.
inputs:
- URL of the Wikipedia page
steps:
- 'Step 1: Import necessary modules from _bootstrap (NIM_MODEL, require_nim_api_key)
and blacknode'
- 'Step 2: Require NVIDIA NIM API key using `require_nim_api_key()`'
- 'Step 3: Create a Graph instance `g`'
- 'Step 4: Add a Literal node for the URL of the Wikipedia page (`url`)'
- 'Step 5: Add an HTTPGet node to fetch the content from the URL (`fetcher`) and connect
it to the Literal node'
- 'Step 6: Add an LLMAgent node with system prompt ''You are a technical writer. Summarise
the text in 3 bullet points.'' and model NIM_MODEL (`summarise`), connecting its
input to the output of `fetcher`'
- 'Step 7: Add a FileWrite node to write the summary to a file named ''summary.txt''
(`writer`), connecting its input to the output of `summarise`'
- 'Step 8: Cook the graph starting from the writer node and print the path where the
summary is written'
- 'Step 1: Import necessary modules from `_bootstrap` and `blacknode`: Run `from _bootstrap
import NIM_MODEL, require_nim_api_key; import blacknode as bn`'
- 'Step 2: Require NVIDIA NIM API key: Call `require_nim_api_key()` to ensure the
API key is set.'
- 'Step 3: Create a graph instance: Initialize `g = bn.Graph()`.'
- 'Step 4: Add nodes for URL, HTTPGet, summarisation, and file writing: `url = g.node(''Literal'',
value=''URL of the Wikipedia page''); fetcher = g.node(''HTTPGet''); summarise =
g.node(''LLMAgent'', system=''You are a technical writer. Summarise the text in
3 bullet points.'', model=NIM_MODEL); writer = g.node(''FileWrite'', path=''summary.txt'')`'
- 'Step 5: Connect nodes with edges: `url.out(''value'') >> fetcher.inp(''url'');
fetcher.out(''text'') >> summarise.inp(''prompt''); summarise.out(''text'') >> writer.inp(''text'')`'
- 'Step 6: Cook the graph to execute and get output: `result = g.cook(writer, ''path'');
print(f''Summary written to: {result}'')`'
outputs:
- Path to the summary file
- Path of the summary file
tags: []
metadata:
source_repo: https://github.com/temiroff/Blacknode.git
@@ -31,7 +30,7 @@ metadata:
# research-pipeline
Fetch a Wikipedia page, summarise it using an AI agent, and write the summary to a file.
Fetch a Wikipedia page, summarise its content using an AI agent, and write the summary to a file.
## Setup
@@ -43,8 +42,8 @@ pip install anthropic>=0.25 docker>=7.1 openai>=1.0
**Setup steps:**
1. Ensure NVIDIA NIM API key is set in the environment or editor UI
1. Install required dependencies using `pip install -r requirements.txt`
1. Ensure NVIDIA NIM API key is set in the environment or editor
1. Install required dependencies: `pip install -r requirements.txt`
## Key Files
@@ -52,39 +51,25 @@ pip install anthropic>=0.25 docker>=7.1 openai>=1.0
## Steps
1. Step 1: Import necessary modules from _bootstrap (NIM_MODEL, require_nim_api_key) and blacknode
2. Step 2: Require NVIDIA NIM API key using `require_nim_api_key()`
3. Step 3: Create a Graph instance `g`
4. Step 4: Add a Literal node for the URL of the Wikipedia page (`url`)
5. Step 5: Add an HTTPGet node to fetch the content from the URL (`fetcher`) and connect it to the Literal node
6. Step 6: Add an LLMAgent node with system prompt 'You are a technical writer. Summarise the text in 3 bullet points.' and model NIM_MODEL (`summarise`), connecting its input to the output of `fetcher`
7. Step 7: Add a FileWrite node to write the summary to a file named 'summary.txt' (`writer`), connecting its input to the output of `summarise`
8. Step 8: Cook the graph starting from the writer node and print the path where the summary is written
1. Step 1: Import necessary modules from `_bootstrap` and `blacknode`: Run `from _bootstrap import NIM_MODEL, require_nim_api_key; import blacknode as bn`
2. Step 2: Require NVIDIA NIM API key: Call `require_nim_api_key()` to ensure the API key is set.
3. Step 3: Create a graph instance: Initialize `g = bn.Graph()`.
4. Step 4: Add nodes for URL, HTTPGet, summarisation, and file writing: `url = g.node('Literal', value='URL of the Wikipedia page'); fetcher = g.node('HTTPGet'); summarise = g.node('LLMAgent', system='You are a technical writer. Summarise the text in 3 bullet points.', model=NIM_MODEL); writer = g.node('FileWrite', path='summary.txt')`
5. Step 5: Connect nodes with edges: `url.out('value') >> fetcher.inp('url'); fetcher.out('text') >> summarise.inp('prompt'); summarise.out('text') >> writer.inp('text')`
6. Step 6: Cook the graph to execute and get output: `result = g.cook(writer, 'path'); print(f'Summary written to: {result}')`
## Implementation Details
```python
from _bootstrap import NIM_MODEL, require_nim_api_key
import blacknode as bn
from _bootstrap import NIM_MODEL, require_nim_api_key; import blacknode as bn
```
```python
g = bn.Graph()
url = g.node('Literal', value='https://en.wikipedia.org/w/api.php?action=query&prop=extracts&exintro=1&explaintext=1&titles=Houdini_(software)&format=json&formatversion=2&origin=*')
fetcher = g.node('HTTPGet')
summarise = g.node('LLMAgent', system='You are a technical writer. Summarise the text in 3 bullet points.', model=NIM_MODEL)
writer = g.node('FileWrite', path='summary.txt')
url = g.node('Literal', value='https://en.wikipedia.org/w/api.php?action=query&prop=extracts&exintro=1&explaintext=1&titles=Houdini_(software)&format=json&formatversion=2&origin=*'); fetcher = g.node('HTTPGet'); summarise = g.node('LLMAgent', system='You are a technical writer. Summarise the text in 3 bullet points.', model=NIM_MODEL); writer = g.node('FileWrite', path='summary.txt')
```
```python
url.out('value') >> fetcher.inp('url')
fetcher.out('text') >> summarise.inp('prompt')
summarise.out('text') >> writer.inp('text')
```
```python
result = g.cook(writer, 'path')
print(f'Summary written to: {result}')
url.out('value') >> fetcher.inp('url'); fetcher.out('text') >> summarise.inp('prompt'); summarise.out('text') >> writer.inp('text')
```
## Inputs
@@ -93,11 +78,11 @@ print(f'Summary written to: {result}')
## Outputs
- Path to the summary file
- Path of the summary file
## Failure Modes
- If the URL is invalid, HTTPGet will fail; if NIM API key is missing, LLMAgent will not function properly
- If the URL is invalid or unreachable, the HTTPGet node will fail; if the summarisation fails, the output text might be empty
## Source
+2 -2
View File
@@ -5,6 +5,6 @@
```python
# How to use this skill
# Inputs: URL of the Wikipedia page
# Process: Step 1: Import necessary modules from _bootstrap (NIM_MODEL, require_nim_api_key) and blacknode → Step 2: Require NVIDIA NIM API key using `require_nim_api_key()` → Step 3: Create a Graph instance `g`
# Outputs: Path to the summary file
# Process: Step 1: Import necessary modules from `_bootstrap` and `blacknode`: Run `from _bootstrap import NIM_MODEL, require_nim_api_key; import blacknode as bn` → Step 2: Require NVIDIA NIM API key: Call `require_nim_api_key()` to ensure the API key is set. → Step 3: Create a graph instance: Initialize `g = bn.Graph()`.
# Outputs: Path of the summary file
```
+10 -12
View File
@@ -1,28 +1,26 @@
{
"name": "research-pipeline",
"version": "1.0.0",
"goal": "Fetch a Wikipedia page, summarise it using an AI agent, and write the summary to a file.",
"goal": "Fetch a Wikipedia page, summarise its content using an AI agent, and write the summary to a file.",
"inputs": [
"URL of the Wikipedia page"
],
"steps": [
"Step 1: Import necessary modules from _bootstrap (NIM_MODEL, require_nim_api_key) and blacknode",
"Step 2: Require NVIDIA NIM API key using `require_nim_api_key()`",
"Step 3: Create a Graph instance `g`",
"Step 4: Add a Literal node for the URL of the Wikipedia page (`url`)",
"Step 5: Add an HTTPGet node to fetch the content from the URL (`fetcher`) and connect it to the Literal node",
"Step 6: Add an LLMAgent node with system prompt 'You are a technical writer. Summarise the text in 3 bullet points.' and model NIM_MODEL (`summarise`), connecting its input to the output of `fetcher`",
"Step 7: Add a FileWrite node to write the summary to a file named 'summary.txt' (`writer`), connecting its input to the output of `summarise`",
"Step 8: Cook the graph starting from the writer node and print the path where the summary is written"
"Step 1: Import necessary modules from `_bootstrap` and `blacknode`: Run `from _bootstrap import NIM_MODEL, require_nim_api_key; import blacknode as bn`",
"Step 2: Require NVIDIA NIM API key: Call `require_nim_api_key()` to ensure the API key is set.",
"Step 3: Create a graph instance: Initialize `g = bn.Graph()`.",
"Step 4: Add nodes for URL, HTTPGet, summarisation, and file writing: `url = g.node('Literal', value='URL of the Wikipedia page'); fetcher = g.node('HTTPGet'); summarise = g.node('LLMAgent', system='You are a technical writer. Summarise the text in 3 bullet points.', model=NIM_MODEL); writer = g.node('FileWrite', path='summary.txt')`",
"Step 5: Connect nodes with edges: `url.out('value') >> fetcher.inp('url'); fetcher.out('text') >> summarise.inp('prompt'); summarise.out('text') >> writer.inp('text')`",
"Step 6: Cook the graph to execute and get output: `result = g.cook(writer, 'path'); print(f'Summary written to: {result}')`"
],
"outputs": [
"Path to the summary file"
"Path of the summary file"
],
"failure_modes": [
"If the URL is invalid, HTTPGet will fail; if NIM API key is missing, LLMAgent will not function properly"
"If the URL is invalid or unreachable, the HTTPGet node will fail; if the summarisation fails, the output text might be empty"
],
"confidence": 0.95,
"explanation": "This workflow can be adapted to fetch and summarise any text from a URL using an AI agent and save the summary to a file.",
"explanation": "This workflow can be adapted to fetch and summarise any Wikipedia page or similar content source.",
"source_repo": "https://github.com/temiroff/Blacknode.git",
"score": 1.0
}
@@ -0,0 +1,94 @@
---
name: three-tier-evaluation-pipeline
version: 1.0.0
description: Run tasks through three evaluation tiers (Run, Trace, Thread) to produce
comprehensive reports with human-in-the-loop validation
inputs:
- query/input text for the task
- search results (for trace tier evaluation)
- evaluation criteria and thresholds
steps:
- 'Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph
engine) to generate initial outputs and results'
- 'Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined
criteria, generating detailed analysis and scoring'
- 'Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion,
approval, and iterative refinement of the output'
outputs:
- Final consolidated report combining results from all three tiers
- Detailed scores and metrics per tier
- Threaded discussion logs for human review and approval
tags: []
metadata:
source_repo: https://github.com/itszhaoziyan-n/AgentKit.git
extracted_at: ''
confidence: 0.95
---
# three-tier-evaluation-pipeline
Run tasks through three evaluation tiers (Run, Trace, Thread) to produce comprehensive reports with human-in-the-loop validation
## Setup
**Dependencies:**
```text
pip install langgraph>=0.3 langchain-core>=0.3 langchain-anthropic>=0.3 langfuse>=2.0 mcp[server]>=1.24 tenacity>=9.0 fastapi>=0.115 psycopg[binary]>=3.1
```
**Setup steps:**
1. Install dependencies with pip install -e .[dev]
1. Start infrastructure: docker compose up -d (PostgreSQL, Langfuse, MCP server)
1. Configure environment variables (DATABASE_URL, MCP_API_KEY, etc.)
1. Run the pipeline: python -m eval.runner --tiers run,thread,trace
## Key Files
- `eval/ - contains the three-tier evaluation logic`
- `scripts/ci_gate.py - threshold update and benchmark validation`
- `agentkit/runtime/ - LangGraph engine for state management and graph execution`
## Steps
1. Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results
2. Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring
3. Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output
## Implementation Details
```python
The eval/ directory implements Run, Trace, and Thread stages with configurable tiers
```
```python
Benchmark suite (40 test cases) validates the pipeline's reliability
```
```python
CI/CD workflows (ci.yml, eval-fast.yml, eval-trace.yml) orchestrate the evaluation pipeline
```
## Inputs
- query/input text for the task
- search results (for trace tier evaluation)
- evaluation criteria and thresholds
## Outputs
- Final consolidated report combining results from all three tiers
- Detailed scores and metrics per tier
- Threaded discussion logs for human review and approval
## Failure Modes
- If the Run tier fails (e.g., code execution error), the pipeline can retry but may produce incomplete outputs
- If the Trace tier LLM-as-Judge produces low-quality evaluations, the Thread tier may need additional human intervention
- Threshold mismatches between tiers could cause the pipeline to exit early or require manual adjustment
## Source
Extracted from: [https://github.com/itszhaoziyan-n/AgentKit.git](https://github.com/itszhaoziyan-n/AgentKit.git)
Confidence: 0.95
@@ -0,0 +1,6 @@
# Commands: three-tier-evaluation-pipeline
## Available Commands
- `/skill three-tier-evaluation-pipeline` — Load this skill
- `/run three-tier-evaluation-pipeline` — Execute workflow
@@ -0,0 +1,10 @@
# Examples: three-tier-evaluation-pipeline
## Usage Example
```python
# How to use this skill
# Inputs: query/input text for the task, search results (for trace tier evaluation), evaluation criteria and thresholds
# Process: Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results → Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring → Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output
# Outputs: Final consolidated report combining results from all three tiers, Detailed scores and metrics per tier, Threaded discussion logs for human review and approval
```
@@ -0,0 +1,29 @@
{
"name": "three-tier-evaluation-pipeline",
"version": "1.0.0",
"goal": "Run tasks through three evaluation tiers (Run, Trace, Thread) to produce comprehensive reports with human-in-the-loop validation",
"inputs": [
"query/input text for the task",
"search results (for trace tier evaluation)",
"evaluation criteria and thresholds"
],
"steps": [
"Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results",
"Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring",
"Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output"
],
"outputs": [
"Final consolidated report combining results from all three tiers",
"Detailed scores and metrics per tier",
"Threaded discussion logs for human review and approval"
],
"failure_modes": [
"If the Run tier fails (e.g., code execution error), the pipeline can retry but may produce incomplete outputs",
"If the Trace tier LLM-as-Judge produces low-quality evaluations, the Thread tier may need additional human intervention",
"Threshold mismatches between tiers could cause the pipeline to exit early or require manual adjustment"
],
"confidence": 0.95,
"explanation": "The AgentKit repository contains a production-ready three-tier evaluation pipeline (Run \u2192 Trace \u2192 Thread) that can be adapted to any task requiring multi-stage validation. This workflow uses LangGraph for orchestration and LangChain for tool integration, making it portable across different agent engineering scenarios. The pattern is reusable because it separates concerns into distinct stages with clear inputs/outputs, allowing teams to plug in different evaluation criteria or human reviewers as needed.",
"source_repo": "https://github.com/itszhaoziyan-n/AgentKit.git",
"score": 1.0
}
@@ -0,0 +1,9 @@
# Tests: three-tier-evaluation-pipeline
## Test Checklist
- [ ] Workflow has at least 3 steps
- [ ] All inputs are defined
- [ ] All outputs are defined
- [ ] Failure modes are documented
- [ ] Skill can be loaded without errors