Compare commits
3 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 1ba56fd7e3 | |||
| d81ddeda88 | |||
| a14f09bec2 |
@@ -34,8 +34,8 @@ scout:
|
|||||||
- 'rag agent workflow'
|
- 'rag agent workflow'
|
||||||
- 'tool calling workflow'
|
- 'tool calling workflow'
|
||||||
filters:
|
filters:
|
||||||
stars_min: 15
|
stars_min: 10
|
||||||
pushed_after: 2026-05-01
|
pushed_after: 2026-02-01
|
||||||
language: Python
|
language: Python
|
||||||
archived: false
|
archived: false
|
||||||
size_max_kb: 10000
|
size_max_kb: 10000
|
||||||
|
|||||||
@@ -0,0 +1,84 @@
|
|||||||
|
---
|
||||||
|
name: agent-supervisor
|
||||||
|
version: 1.0.0
|
||||||
|
description: Demonstrate a supervisor-worker architecture for intelligent task delegation
|
||||||
|
and real-time decision-making.
|
||||||
|
inputs:
|
||||||
|
- name: OPENAI_API_KEY
|
||||||
|
description: OpenAI API key for language models.
|
||||||
|
- name: TAVILY_API_KEY
|
||||||
|
description: Tavily API key for search functionality.
|
||||||
|
steps:
|
||||||
|
- step: 1
|
||||||
|
action: Load environment variables.
|
||||||
|
details: Set the OPENAI_API_KEY and TAVILY_API_KEY environment variables.
|
||||||
|
- step: 2
|
||||||
|
action: Configure LangChain tools.
|
||||||
|
details: Initialize TavilySearchResults and PythonREPLTool.
|
||||||
|
- step: 3
|
||||||
|
action: Define agent nodes.
|
||||||
|
details: Create functions for the Researcher and Coder agents that process state
|
||||||
|
through their respective tasks.
|
||||||
|
- step: 4
|
||||||
|
action: Set up supervisor agent.
|
||||||
|
details: Create a supervisor agent function that decides which worker should act
|
||||||
|
next based on user input.
|
||||||
|
- step: 5
|
||||||
|
action: Build state graph.
|
||||||
|
details: Construct the state graph with nodes for each agent and edges connecting
|
||||||
|
them to the supervisor node.
|
||||||
|
- step: 6
|
||||||
|
action: Add conditional edges.
|
||||||
|
details: Define conditions for transitioning between agents based on their responses.
|
||||||
|
- step: 7
|
||||||
|
action: Compile graph.
|
||||||
|
details: Compile the state graph into a runnable workflow.
|
||||||
|
- step: 8
|
||||||
|
action: Run example queries.
|
||||||
|
details: Stream through the workflow with example inputs to demonstrate its functionality.
|
||||||
|
outputs:
|
||||||
|
- name: 'Example 1: Code Hello World'
|
||||||
|
description: A demonstration of coding a simple hello world program.
|
||||||
|
- name: 'Example 2: Research Report'
|
||||||
|
description: A demonstration of researching and writing a brief report on pikas.
|
||||||
|
tags: []
|
||||||
|
metadata:
|
||||||
|
source_repo: https://github.com/extrawest/multi_agent_workflow_demo_in_langgraph.git
|
||||||
|
extracted_at: ''
|
||||||
|
confidence: 0.9
|
||||||
|
---
|
||||||
|
|
||||||
|
# agent-supervisor
|
||||||
|
|
||||||
|
Demonstrate a supervisor-worker architecture for intelligent task delegation and real-time decision-making.
|
||||||
|
|
||||||
|
## Steps
|
||||||
|
|
||||||
|
1. {'step': 1, 'action': 'Load environment variables.', 'details': 'Set the OPENAI_API_KEY and TAVILY_API_KEY environment variables.'}
|
||||||
|
2. {'step': 2, 'action': 'Configure LangChain tools.', 'details': 'Initialize TavilySearchResults and PythonREPLTool.'}
|
||||||
|
3. {'step': 3, 'action': 'Define agent nodes.', 'details': 'Create functions for the Researcher and Coder agents that process state through their respective tasks.'}
|
||||||
|
4. {'step': 4, 'action': 'Set up supervisor agent.', 'details': 'Create a supervisor agent function that decides which worker should act next based on user input.'}
|
||||||
|
5. {'step': 5, 'action': 'Build state graph.', 'details': 'Construct the state graph with nodes for each agent and edges connecting them to the supervisor node.'}
|
||||||
|
6. {'step': 6, 'action': 'Add conditional edges.', 'details': 'Define conditions for transitioning between agents based on their responses.'}
|
||||||
|
7. {'step': 7, 'action': 'Compile graph.', 'details': 'Compile the state graph into a runnable workflow.'}
|
||||||
|
8. {'step': 8, 'action': 'Run example queries.', 'details': 'Stream through the workflow with example inputs to demonstrate its functionality.'}
|
||||||
|
|
||||||
|
## Inputs
|
||||||
|
|
||||||
|
- {'name': 'OPENAI_API_KEY', 'description': 'OpenAI API key for language models.'}
|
||||||
|
- {'name': 'TAVILY_API_KEY', 'description': 'Tavily API key for search functionality.'}
|
||||||
|
|
||||||
|
## Outputs
|
||||||
|
|
||||||
|
- {'name': 'Example 1: Code Hello World', 'description': 'A demonstration of coding a simple hello world program.'}
|
||||||
|
- {'name': 'Example 2: Research Report', 'description': 'A demonstration of researching and writing a brief report on pikas.'}
|
||||||
|
|
||||||
|
## Failure Modes
|
||||||
|
|
||||||
|
- {'mode': 'Invalid API keys', 'description': 'The workflow may fail if the provided API keys are invalid or expired.'}
|
||||||
|
- {'mode': 'Insufficient permissions', 'description': 'The workflow may fail if the user does not have sufficient permissions to use the Tavily search functionality.'}
|
||||||
|
|
||||||
|
## Source
|
||||||
|
|
||||||
|
Extracted from: [https://github.com/extrawest/multi_agent_workflow_demo_in_langgraph.git](https://github.com/extrawest/multi_agent_workflow_demo_in_langgraph.git)
|
||||||
|
Confidence: 0.9
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
# Commands: agent-supervisor
|
||||||
|
|
||||||
|
## Available Commands
|
||||||
|
|
||||||
|
- `/skill agent-supervisor` — Load this skill
|
||||||
|
- `/run agent-supervisor` — Execute workflow
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
# Examples: agent-supervisor
|
||||||
|
|
||||||
|
## Usage Example
|
||||||
|
|
||||||
|
```python
|
||||||
|
# How to use this skill
|
||||||
|
# Inputs: {'name': 'OPENAI_API_KEY', 'description': 'OpenAI API key for language models.'}, {'name': 'TAVILY_API_KEY', 'description': 'Tavily API key for search functionality.'}
|
||||||
|
# Process: {'step': 1, 'action': 'Load environment variables.', 'details': 'Set the OPENAI_API_KEY and TAVILY_API_KEY environment variables.'} → {'step': 2, 'action': 'Configure LangChain tools.', 'details': 'Initialize TavilySearchResults and PythonREPLTool.'} → {'step': 3, 'action': 'Define agent nodes.', 'details': 'Create functions for the Researcher and Coder agents that process state through their respective tasks.'}
|
||||||
|
# Outputs: {'name': 'Example 1: Code Hello World', 'description': 'A demonstration of coding a simple hello world program.'}, {'name': 'Example 2: Research Report', 'description': 'A demonstration of researching and writing a brief report on pikas.'}
|
||||||
|
```
|
||||||
@@ -0,0 +1,81 @@
|
|||||||
|
{
|
||||||
|
"name": "agent-supervisor",
|
||||||
|
"version": "1.0.0",
|
||||||
|
"goal": "Demonstrate a supervisor-worker architecture for intelligent task delegation and real-time decision-making.",
|
||||||
|
"inputs": [
|
||||||
|
{
|
||||||
|
"name": "OPENAI_API_KEY",
|
||||||
|
"description": "OpenAI API key for language models."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "TAVILY_API_KEY",
|
||||||
|
"description": "Tavily API key for search functionality."
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"steps": [
|
||||||
|
{
|
||||||
|
"step": 1,
|
||||||
|
"action": "Load environment variables.",
|
||||||
|
"details": "Set the OPENAI_API_KEY and TAVILY_API_KEY environment variables."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"step": 2,
|
||||||
|
"action": "Configure LangChain tools.",
|
||||||
|
"details": "Initialize TavilySearchResults and PythonREPLTool."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"step": 3,
|
||||||
|
"action": "Define agent nodes.",
|
||||||
|
"details": "Create functions for the Researcher and Coder agents that process state through their respective tasks."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"step": 4,
|
||||||
|
"action": "Set up supervisor agent.",
|
||||||
|
"details": "Create a supervisor agent function that decides which worker should act next based on user input."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"step": 5,
|
||||||
|
"action": "Build state graph.",
|
||||||
|
"details": "Construct the state graph with nodes for each agent and edges connecting them to the supervisor node."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"step": 6,
|
||||||
|
"action": "Add conditional edges.",
|
||||||
|
"details": "Define conditions for transitioning between agents based on their responses."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"step": 7,
|
||||||
|
"action": "Compile graph.",
|
||||||
|
"details": "Compile the state graph into a runnable workflow."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"step": 8,
|
||||||
|
"action": "Run example queries.",
|
||||||
|
"details": "Stream through the workflow with example inputs to demonstrate its functionality."
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "Example 1: Code Hello World",
|
||||||
|
"description": "A demonstration of coding a simple hello world program."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "Example 2: Research Report",
|
||||||
|
"description": "A demonstration of researching and writing a brief report on pikas."
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"failure_modes": [
|
||||||
|
{
|
||||||
|
"mode": "Invalid API keys",
|
||||||
|
"description": "The workflow may fail if the provided API keys are invalid or expired."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"mode": "Insufficient permissions",
|
||||||
|
"description": "The workflow may fail if the user does not have sufficient permissions to use the Tavily search functionality."
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"confidence": 0.9,
|
||||||
|
"explanation": "This workflow demonstrates a hierarchical multi-agent system where a supervisor agent makes routing decisions based on user input, delegating tasks to specialized worker agents (Researcher and Coder). It is designed to be reusable for similar task delegation scenarios.",
|
||||||
|
"source_repo": "https://github.com/extrawest/multi_agent_workflow_demo_in_langgraph.git",
|
||||||
|
"score": 1.0
|
||||||
|
}
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
# Tests: agent-supervisor
|
||||||
|
|
||||||
|
## Test Checklist
|
||||||
|
|
||||||
|
- [ ] Workflow has at least 3 steps
|
||||||
|
- [ ] All inputs are defined
|
||||||
|
- [ ] All outputs are defined
|
||||||
|
- [ ] Failure modes are documented
|
||||||
|
- [ ] Skill can be loaded without errors
|
||||||
@@ -0,0 +1,94 @@
|
|||||||
|
---
|
||||||
|
name: three-tier-evaluation-pipeline
|
||||||
|
version: 1.0.0
|
||||||
|
description: Run tasks through three evaluation tiers (Run, Trace, Thread) to produce
|
||||||
|
comprehensive reports with human-in-the-loop validation
|
||||||
|
inputs:
|
||||||
|
- query/input text for the task
|
||||||
|
- search results (for trace tier evaluation)
|
||||||
|
- evaluation criteria and thresholds
|
||||||
|
steps:
|
||||||
|
- 'Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph
|
||||||
|
engine) to generate initial outputs and results'
|
||||||
|
- 'Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined
|
||||||
|
criteria, generating detailed analysis and scoring'
|
||||||
|
- 'Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion,
|
||||||
|
approval, and iterative refinement of the output'
|
||||||
|
outputs:
|
||||||
|
- Final consolidated report combining results from all three tiers
|
||||||
|
- Detailed scores and metrics per tier
|
||||||
|
- Threaded discussion logs for human review and approval
|
||||||
|
tags: []
|
||||||
|
metadata:
|
||||||
|
source_repo: https://github.com/itszhaoziyan-n/AgentKit.git
|
||||||
|
extracted_at: ''
|
||||||
|
confidence: 0.95
|
||||||
|
---
|
||||||
|
|
||||||
|
# three-tier-evaluation-pipeline
|
||||||
|
|
||||||
|
Run tasks through three evaluation tiers (Run, Trace, Thread) to produce comprehensive reports with human-in-the-loop validation
|
||||||
|
|
||||||
|
## Setup
|
||||||
|
|
||||||
|
**Dependencies:**
|
||||||
|
|
||||||
|
```text
|
||||||
|
pip install langgraph>=0.3 langchain-core>=0.3 langchain-anthropic>=0.3 langfuse>=2.0 mcp[server]>=1.24 tenacity>=9.0 fastapi>=0.115 psycopg[binary]>=3.1
|
||||||
|
```
|
||||||
|
|
||||||
|
**Setup steps:**
|
||||||
|
|
||||||
|
1. Install dependencies with pip install -e .[dev]
|
||||||
|
1. Start infrastructure: docker compose up -d (PostgreSQL, Langfuse, MCP server)
|
||||||
|
1. Configure environment variables (DATABASE_URL, MCP_API_KEY, etc.)
|
||||||
|
1. Run the pipeline: python -m eval.runner --tiers run,thread,trace
|
||||||
|
|
||||||
|
## Key Files
|
||||||
|
|
||||||
|
- `eval/ - contains the three-tier evaluation logic`
|
||||||
|
- `scripts/ci_gate.py - threshold update and benchmark validation`
|
||||||
|
- `agentkit/runtime/ - LangGraph engine for state management and graph execution`
|
||||||
|
|
||||||
|
## Steps
|
||||||
|
|
||||||
|
1. Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results
|
||||||
|
2. Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring
|
||||||
|
3. Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output
|
||||||
|
|
||||||
|
## Implementation Details
|
||||||
|
|
||||||
|
```python
|
||||||
|
The eval/ directory implements Run, Trace, and Thread stages with configurable tiers
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
Benchmark suite (40 test cases) validates the pipeline's reliability
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
CI/CD workflows (ci.yml, eval-fast.yml, eval-trace.yml) orchestrate the evaluation pipeline
|
||||||
|
```
|
||||||
|
|
||||||
|
## Inputs
|
||||||
|
|
||||||
|
- query/input text for the task
|
||||||
|
- search results (for trace tier evaluation)
|
||||||
|
- evaluation criteria and thresholds
|
||||||
|
|
||||||
|
## Outputs
|
||||||
|
|
||||||
|
- Final consolidated report combining results from all three tiers
|
||||||
|
- Detailed scores and metrics per tier
|
||||||
|
- Threaded discussion logs for human review and approval
|
||||||
|
|
||||||
|
## Failure Modes
|
||||||
|
|
||||||
|
- If the Run tier fails (e.g., code execution error), the pipeline can retry but may produce incomplete outputs
|
||||||
|
- If the Trace tier LLM-as-Judge produces low-quality evaluations, the Thread tier may need additional human intervention
|
||||||
|
- Threshold mismatches between tiers could cause the pipeline to exit early or require manual adjustment
|
||||||
|
|
||||||
|
## Source
|
||||||
|
|
||||||
|
Extracted from: [https://github.com/itszhaoziyan-n/AgentKit.git](https://github.com/itszhaoziyan-n/AgentKit.git)
|
||||||
|
Confidence: 0.95
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
# Commands: three-tier-evaluation-pipeline
|
||||||
|
|
||||||
|
## Available Commands
|
||||||
|
|
||||||
|
- `/skill three-tier-evaluation-pipeline` — Load this skill
|
||||||
|
- `/run three-tier-evaluation-pipeline` — Execute workflow
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
# Examples: three-tier-evaluation-pipeline
|
||||||
|
|
||||||
|
## Usage Example
|
||||||
|
|
||||||
|
```python
|
||||||
|
# How to use this skill
|
||||||
|
# Inputs: query/input text for the task, search results (for trace tier evaluation), evaluation criteria and thresholds
|
||||||
|
# Process: Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results → Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring → Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output
|
||||||
|
# Outputs: Final consolidated report combining results from all three tiers, Detailed scores and metrics per tier, Threaded discussion logs for human review and approval
|
||||||
|
```
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
{
|
||||||
|
"name": "three-tier-evaluation-pipeline",
|
||||||
|
"version": "1.0.0",
|
||||||
|
"goal": "Run tasks through three evaluation tiers (Run, Trace, Thread) to produce comprehensive reports with human-in-the-loop validation",
|
||||||
|
"inputs": [
|
||||||
|
"query/input text for the task",
|
||||||
|
"search results (for trace tier evaluation)",
|
||||||
|
"evaluation criteria and thresholds"
|
||||||
|
],
|
||||||
|
"steps": [
|
||||||
|
"Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results",
|
||||||
|
"Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring",
|
||||||
|
"Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output"
|
||||||
|
],
|
||||||
|
"outputs": [
|
||||||
|
"Final consolidated report combining results from all three tiers",
|
||||||
|
"Detailed scores and metrics per tier",
|
||||||
|
"Threaded discussion logs for human review and approval"
|
||||||
|
],
|
||||||
|
"failure_modes": [
|
||||||
|
"If the Run tier fails (e.g., code execution error), the pipeline can retry but may produce incomplete outputs",
|
||||||
|
"If the Trace tier LLM-as-Judge produces low-quality evaluations, the Thread tier may need additional human intervention",
|
||||||
|
"Threshold mismatches between tiers could cause the pipeline to exit early or require manual adjustment"
|
||||||
|
],
|
||||||
|
"confidence": 0.95,
|
||||||
|
"explanation": "The AgentKit repository contains a production-ready three-tier evaluation pipeline (Run \u2192 Trace \u2192 Thread) that can be adapted to any task requiring multi-stage validation. This workflow uses LangGraph for orchestration and LangChain for tool integration, making it portable across different agent engineering scenarios. The pattern is reusable because it separates concerns into distinct stages with clear inputs/outputs, allowing teams to plug in different evaluation criteria or human reviewers as needed.",
|
||||||
|
"source_repo": "https://github.com/itszhaoziyan-n/AgentKit.git",
|
||||||
|
"score": 1.0
|
||||||
|
}
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
# Tests: three-tier-evaluation-pipeline
|
||||||
|
|
||||||
|
## Test Checklist
|
||||||
|
|
||||||
|
- [ ] Workflow has at least 3 steps
|
||||||
|
- [ ] All inputs are defined
|
||||||
|
- [ ] All outputs are defined
|
||||||
|
- [ ] Failure modes are documented
|
||||||
|
- [ ] Skill can be loaded without errors
|
||||||
Reference in New Issue
Block a user