Compare commits

..

1 Commits

Author SHA1 Message Date
Epictetus 271f79610d Add 5 skills from LFM + 12 skills total
New skills:
- blacknode-graph-workflow
- multi-agent-workflow-execution
- langgraph-agent-workflow
- langgraph-multi-agent-router
- three-tier-evaluation-pipeline

Config: LLM pipeline uses LFM on llama.cpp (8080)
2026-08-05 17:05:21 +00:00
21 changed files with 650 additions and 3 deletions
+3 -3
View File
@@ -16,10 +16,10 @@ llm:
api_key: ""
max_tokens: 8000
# Secondary LLM for pipeline tasks — uses Ollama on 3060 (non-reasoning model)
# Secondary LLM for pipeline tasks — uses LFM on 3060 (llama.cpp)
llm_pipeline:
base_url: http://100.64.0.4:11434
model: qwen2.5:7b
base_url: http://100.64.0.4:8080
model: C:\models\LFM2.5-2.6B-Q4_K_M.gguf
api_key: ""
max_tokens: 6000
+115
View File
@@ -0,0 +1,115 @@
---
name: langgraph-agent-workflow
version: 1.0.0
description: Orchestrate multi-step AI agents using LangGraph with SerperDevTool for
RAG, code execution, and citation generation
inputs:
- LangGraph chain configuration files defining agent workflows
- SerperDevTool integration for LLM tool access
- React agent creation scripts via create_react_agent
- Knowledge graph retrieval and citation generation pipelines
steps:
- 'Step 1: Define LangGraph chain architecture with SerperDevTool integration - Create
a LangGraph chain that combines retrieval, reasoning, and response generation using
SerperDevTool for tool access'
- 'Step 2: Implement React agent wrapper - Use create_react_agent to build a frontend
agent that can interact with the LangGraph chain'
- 'Step 3: Configure RAG pipeline - Set up vector search (Qdrant/OpenSearch) and knowledge
graph retrieval (Neo4j/ArangoDB) with citation generation'
- 'Step 4: Add code execution sandbox - Integrate artifact generation capabilities
for code-related tasks'
- "Step 5: Orchestrate multi-step research workflow - Chain search \u2192 deep research\
\ \u2192 agent response in a single LangGraph workflow"
outputs:
- Reusable LangGraph chain definition (pyfile) with configurable steps
- React agent frontend component that can be deployed independently
- RAG pipeline that generates block citations and grounded answers
- Code execution sandbox for artifact generation
- Documentation for parameterizing workflows for different tasks
tags: []
metadata:
source_repo: https://github.com/pipeshub-ai/pipeshub-ai.git
extracted_at: ''
confidence: 0.95
---
# langgraph-agent-workflow
Orchestrate multi-step AI agents using LangGraph with SerperDevTool for RAG, code execution, and citation generation
## Setup
**Dependencies:**
```text
pip install langgraph>=0.7.0 serper-dev-tool>=0.1.0 qdrant-client or opensearch-dsl neo4j-driver or arango-database-driver react, next.js
```
**Setup steps:**
1. Install LangGraph and SerperDevTool dependencies
1. Configure vector store (Qdrant/OpenSearch) and knowledge graph (Neo4j/ArangoDB)
1. Define chain topology with retrieval, reasoning, and response steps
1. Build React agent frontend using create_react_agent
1. Test multi-step agent workflows end-to-end
## Key Files
- `pipeshub-ai/workflows/agent_chain.py - Main LangGraph chain definition`
- `pipeshub-ai/workflows/agent_react.py - React agent wrapper`
- `pipeshub-ai/workflows/rag_pipeline.py - RAG with citation generation`
- `pipeshub-ai/workflows/code_sandbox.py - Code execution sandbox`
## Steps
1. Step 1: Define LangGraph chain architecture with SerperDevTool integration - Create a LangGraph chain that combines retrieval, reasoning, and response generation using SerperDevTool for tool access
2. Step 2: Implement React agent wrapper - Use create_react_agent to build a frontend agent that can interact with the LangGraph chain
3. Step 3: Configure RAG pipeline - Set up vector search (Qdrant/OpenSearch) and knowledge graph retrieval (Neo4j/ArangoDB) with citation generation
4. Step 4: Add code execution sandbox - Integrate artifact generation capabilities for code-related tasks
5. Step 5: Orchestrate multi-step research workflow - Chain search → deep research → agent response in a single LangGraph workflow
## Implementation Details
```python
chain = LangGraph()
```
```python
chain.add_step(SerperDevToolAgent())
```
```python
agent = create_react_agent(chain, SerperDevToolAgent())
```
```python
workflow = chain.start()
```
## Inputs
- LangGraph chain configuration files defining agent workflows
- SerperDevTool integration for LLM tool access
- React agent creation scripts via create_react_agent
- Knowledge graph retrieval and citation generation pipelines
## Outputs
- Reusable LangGraph chain definition (pyfile) with configurable steps
- React agent frontend component that can be deployed independently
- RAG pipeline that generates block citations and grounded answers
- Code execution sandbox for artifact generation
- Documentation for parameterizing workflows for different tasks
## Failure Modes
- GraphDB connection failures if Neo4j/ArangoDB is not properly configured
- Vector store unavailability (Qdrant/OpenSearch) causing RAG pipeline to fail
- LLM tool access errors if SerperDevTool is not properly initialized
- Agent timeout if complex multi-step reasoning exceeds time limits
- Sandbox execution failures if code has security vulnerabilities or infinite loops
## Source
Extracted from: [https://github.com/pipeshub-ai/pipeshub-ai.git](https://github.com/pipeshub-ai/pipeshub-ai.git)
Confidence: 0.95
@@ -0,0 +1,6 @@
# Commands: langgraph-agent-workflow
## Available Commands
- `/skill langgraph-agent-workflow` — Load this skill
- `/run langgraph-agent-workflow` — Execute workflow
@@ -0,0 +1,10 @@
# Examples: langgraph-agent-workflow
## Usage Example
```python
# How to use this skill
# Inputs: LangGraph chain configuration files defining agent workflows, SerperDevTool integration for LLM tool access, React agent creation scripts via create_react_agent, Knowledge graph retrieval and citation generation pipelines
# Process: Step 1: Define LangGraph chain architecture with SerperDevTool integration - Create a LangGraph chain that combines retrieval, reasoning, and response generation using SerperDevTool for tool access → Step 2: Implement React agent wrapper - Use create_react_agent to build a frontend agent that can interact with the LangGraph chain → Step 3: Configure RAG pipeline - Set up vector search (Qdrant/OpenSearch) and knowledge graph retrieval (Neo4j/ArangoDB) with citation generation
# Outputs: Reusable LangGraph chain definition (pyfile) with configurable steps, React agent frontend component that can be deployed independently, RAG pipeline that generates block citations and grounded answers, Code execution sandbox for artifact generation, Documentation for parameterizing workflows for different tasks
```
@@ -0,0 +1,36 @@
{
"name": "langgraph-agent-workflow",
"version": "1.0.0",
"goal": "Orchestrate multi-step AI agents using LangGraph with SerperDevTool for RAG, code execution, and citation generation",
"inputs": [
"LangGraph chain configuration files defining agent workflows",
"SerperDevTool integration for LLM tool access",
"React agent creation scripts via create_react_agent",
"Knowledge graph retrieval and citation generation pipelines"
],
"steps": [
"Step 1: Define LangGraph chain architecture with SerperDevTool integration - Create a LangGraph chain that combines retrieval, reasoning, and response generation using SerperDevTool for tool access",
"Step 2: Implement React agent wrapper - Use create_react_agent to build a frontend agent that can interact with the LangGraph chain",
"Step 3: Configure RAG pipeline - Set up vector search (Qdrant/OpenSearch) and knowledge graph retrieval (Neo4j/ArangoDB) with citation generation",
"Step 4: Add code execution sandbox - Integrate artifact generation capabilities for code-related tasks",
"Step 5: Orchestrate multi-step research workflow - Chain search \u2192 deep research \u2192 agent response in a single LangGraph workflow"
],
"outputs": [
"Reusable LangGraph chain definition (pyfile) with configurable steps",
"React agent frontend component that can be deployed independently",
"RAG pipeline that generates block citations and grounded answers",
"Code execution sandbox for artifact generation",
"Documentation for parameterizing workflows for different tasks"
],
"failure_modes": [
"GraphDB connection failures if Neo4j/ArangoDB is not properly configured",
"Vector store unavailability (Qdrant/OpenSearch) causing RAG pipeline to fail",
"LLM tool access errors if SerperDevTool is not properly initialized",
"Agent timeout if complex multi-step reasoning exceeds time limits",
"Sandbox execution failures if code has security vulnerabilities or infinite loops"
],
"confidence": 0.95,
"explanation": "PipesHub provides a reusable LangGraph-based agent workflow framework that can be parameterized for different tasks. The core pattern involves defining a LangGraph chain with SerperDevTool integration for tool access, creating a React agent wrapper, and configuring RAG pipelines with citation generation. This framework can be reused across RAG, code execution, and research workflows by adjusting the chain definition and agent configuration.",
"source_repo": "https://github.com/pipeshub-ai/pipeshub-ai.git",
"score": 1.0
}
+9
View File
@@ -0,0 +1,9 @@
# Tests: langgraph-agent-workflow
## Test Checklist
- [ ] Workflow has at least 3 steps
- [ ] All inputs are defined
- [ ] All outputs are defined
- [ ] Failure modes are documented
- [ ] Skill can be loaded without errors
@@ -0,0 +1,96 @@
---
name: langgraph-multi-agent-router
version: 1.0.0
description: Orchestrate a multi-agent workflow where specialized agents collaborate
sequentially to gather information, structure it, and generate a final response
inputs:
- User query string (e.g., destination location)
- BedrockModel configuration (model_id, temperature, top_p)
- Pre-configured agents with specific system prompts and tool sets
steps:
- Researcher agent executes with system prompt to gather raw destination facts (places,
history, accommodations, food, web pages) using BedrockModel and available tools
(calculator, current_time)
- Travel guide agent receives raw research output and structures it into labeled sections
(Must-See Attractions, Historical Highlights, Accommodation Areas, Culinary Delights,
Suggested Web Pages)
- Writer agent receives the structured guide and synthesizes it into a professional
client-facing response with clear formatting and emphasis on the suggested web pages
outputs:
- Raw research data (JSON string containing gathered facts)
- Structured guide content (markdown-formatted travel guide with labeled sections)
- Final client response (professional formatted response ready for delivery)
tags: []
metadata:
source_repo: https://github.com/omerbsezer/Fast-LLM-Agent-MCP.git
extracted_at: ''
confidence: 0.95
---
# langgraph-multi-agent-router
Orchestrate a multi-agent workflow where specialized agents collaborate sequentially to gather information, structure it, and generate a final response
## Setup
**Dependencies:**
```text
pip install langchain langgraph bedrock-model pydantic
```
**Setup steps:**
1. Install langchain and langgraph packages
1. Configure BedrockModel with desired parameters (model_id, temperature, top_p)
1. Create three Agent instances with specific system prompts and tool sets
1. Initialize LangGraph with the agent chain and run the workflow
## Key Files
- `agents/langchain_langgraph/00-basic-agent/agent.py`
- `agents/langchain_langgraph/02-agent-with-tools-structured-output/agent.py`
- `agents/langchain_langgraph/08-langgraph-multi-agents-sequential-pattern/agent.py`
## Steps
1. Researcher agent executes with system prompt to gather raw destination facts (places, history, accommodations, food, web pages) using BedrockModel and available tools (calculator, current_time)
2. Travel guide agent receives raw research output and structures it into labeled sections (Must-See Attractions, Historical Highlights, Accommodation Areas, Culinary Delights, Suggested Web Pages)
3. Writer agent receives the structured guide and synthesizes it into a professional client-facing response with clear formatting and emphasis on the suggested web pages
## Implementation Details
```python
Researcher agent uses BedrockModel with temperature=0.7, top_p=0.9 to gather destination facts
```
```python
Travel guide agent receives raw output and formats into 5 labeled sections
```
```python
Writer agent takes structured guide and writes professional client response
```
## Inputs
- User query string (e.g., destination location)
- BedrockModel configuration (model_id, temperature, top_p)
- Pre-configured agents with specific system prompts and tool sets
## Outputs
- Raw research data (JSON string containing gathered facts)
- Structured guide content (markdown-formatted travel guide with labeled sections)
- Final client response (professional formatted response ready for delivery)
## Failure Modes
- Researcher agent fails to gather sufficient data or returns incomplete results
- Travel guide agent fails to structure information correctly or produces unreadable output
- Writer agent fails to format the final response properly or loses key information from the guide
## Source
Extracted from: [https://github.com/omerbsezer/Fast-LLM-Agent-MCP.git](https://github.com/omerbsezer/Fast-LLM-Agent-MCP.git)
Confidence: 0.95
@@ -0,0 +1,6 @@
# Commands: langgraph-multi-agent-router
## Available Commands
- `/skill langgraph-multi-agent-router` — Load this skill
- `/run langgraph-multi-agent-router` — Execute workflow
@@ -0,0 +1,10 @@
# Examples: langgraph-multi-agent-router
## Usage Example
```python
# How to use this skill
# Inputs: User query string (e.g., destination location), BedrockModel configuration (model_id, temperature, top_p), Pre-configured agents with specific system prompts and tool sets
# Process: Researcher agent executes with system prompt to gather raw destination facts (places, history, accommodations, food, web pages) using BedrockModel and available tools (calculator, current_time) → Travel guide agent receives raw research output and structures it into labeled sections (Must-See Attractions, Historical Highlights, Accommodation Areas, Culinary Delights, Suggested Web Pages) → Writer agent receives the structured guide and synthesizes it into a professional client-facing response with clear formatting and emphasis on the suggested web pages
# Outputs: Raw research data (JSON string containing gathered facts), Structured guide content (markdown-formatted travel guide with labeled sections), Final client response (professional formatted response ready for delivery)
```
@@ -0,0 +1,29 @@
{
"name": "langgraph-multi-agent-router",
"version": "1.0.0",
"goal": "Orchestrate a multi-agent workflow where specialized agents collaborate sequentially to gather information, structure it, and generate a final response",
"inputs": [
"User query string (e.g., destination location)",
"BedrockModel configuration (model_id, temperature, top_p)",
"Pre-configured agents with specific system prompts and tool sets"
],
"steps": [
"Researcher agent executes with system prompt to gather raw destination facts (places, history, accommodations, food, web pages) using BedrockModel and available tools (calculator, current_time)",
"Travel guide agent receives raw research output and structures it into labeled sections (Must-See Attractions, Historical Highlights, Accommodation Areas, Culinary Delights, Suggested Web Pages)",
"Writer agent receives the structured guide and synthesizes it into a professional client-facing response with clear formatting and emphasis on the suggested web pages"
],
"outputs": [
"Raw research data (JSON string containing gathered facts)",
"Structured guide content (markdown-formatted travel guide with labeled sections)",
"Final client response (professional formatted response ready for delivery)"
],
"failure_modes": [
"Researcher agent fails to gather sufficient data or returns incomplete results",
"Travel guide agent fails to structure information correctly or produces unreadable output",
"Writer agent fails to format the final response properly or loses key information from the guide"
],
"confidence": 0.95,
"explanation": "This workflow demonstrates a reusable multi-stage agent pattern where specialized agents collaborate in sequence. The Researcher agent gathers raw information using a domain-specific model, the Travel Guide agent structures that information into a consistent format, and the Writer agent synthesizes the final output. This pattern can be adapted to other domains (e.g., code generation, data analysis, research workflows) by swapping the agent types and system prompts while maintaining the same three-step structure.",
"source_repo": "https://github.com/omerbsezer/Fast-LLM-Agent-MCP.git",
"score": 1.0
}
@@ -0,0 +1,9 @@
# Tests: langgraph-multi-agent-router
## Test Checklist
- [ ] Workflow has at least 3 steps
- [ ] All inputs are defined
- [ ] All outputs are defined
- [ ] Failure modes are documented
- [ ] Skill can be loaded without errors
@@ -0,0 +1,113 @@
---
name: multi-agent-workflow-execution
version: 1.0.0
description: Execute multi-agent AI workflows defined in YAML blueprints by creating
sessions, submitting user prompts, and polling for completion until final answers
are returned.
inputs:
- Blueprint ID or name (to identify the workflow to execute)
- User shortcut (authentication identifier for the user)
- User question or prompt (input to the workflow)
- Base URL of the UnifAI API (endpoint for session management)
- Polling interval (seconds between status checks during execution)
steps:
- Resolve the blueprint ID from either direct ID or name lookup via the API, handling
cases where the blueprint is not found or not unique
- Create a new session from the resolved blueprint using the session creation endpoint
- Submit the session with the user's prompt to start the multi-agent workflow execution
- Poll the session status at regular intervals until the session completes, fails,
or is cancelled
- Retrieve and return the final answer from the completed workflow
outputs:
- Final workflow result or answer (text or structured data)
- Session status (completed, failed, or cancelled)
- Error details if the workflow execution fails or times out
tags: []
metadata:
source_repo: https://github.com/redhat-community-ai-tools/UnifAI.git
extracted_at: ''
confidence: 0.95
---
# multi-agent-workflow-execution
Execute multi-agent AI workflows defined in YAML blueprints by creating sessions, submitting user prompts, and polling for completion until final answers are returned.
## Setup
**Dependencies:**
```text
pip install requests urllib3 python-langgraph temporalio
```
**Setup steps:**
1. Install Python 3.11+ and required packages (requests, langgraph, temporalio)
1. Configure API base URL and user credentials in environment variables or config
1. Define or select a blueprint from the available workflows in the system
1. Run the execution_workflow.py script with blueprint ID/name and user prompt
## Key Files
- `scripts/execution_workflow.py - Main workflow execution script`
- `multi-agent/lib/mas/engine/ - LangGraph-based orchestration modules`
- `multi-agent/lib/mas/elements/ - Node definitions (custom_agent_node, merger_node, etc.)`
- `multi-agent/lib/mas/blueprints/ - Blueprint resolution and validation logic`
## Steps
1. Resolve the blueprint ID from either direct ID or name lookup via the API, handling cases where the blueprint is not found or not unique
2. Create a new session from the resolved blueprint using the session creation endpoint
3. Submit the session with the user's prompt to start the multi-agent workflow execution
4. Poll the session status at regular intervals until the session completes, fails, or is cancelled
5. Retrieve and return the final answer from the completed workflow
## Implementation Details
```python
resolve_blueprint_id() - Resolves blueprint by ID or name lookup with error handling
```
```python
create_session() - Creates a new session from a blueprint via POST /user.session.create
```
```python
submit_session() - Submits user prompt to start workflow via POST /user.session.submit
```
```python
poll_session_status() - Polls session.stream.status at configurable intervals
```
```python
get_final_answer() - Retrieves final output via GET /session.chat.get
```
## Inputs
- Blueprint ID or name (to identify the workflow to execute)
- User shortcut (authentication identifier for the user)
- User question or prompt (input to the workflow)
- Base URL of the UnifAI API (endpoint for session management)
- Polling interval (seconds between status checks during execution)
## Outputs
- Final workflow result or answer (text or structured data)
- Session status (completed, failed, or cancelled)
- Error details if the workflow execution fails or times out
## Failure Modes
- Blueprint not found or not unique - script exits with an error listing available blueprints
- Session creation fails - may be due to invalid blueprint ID, authentication issues, or rate limiting
- Session submission fails - could be due to network issues, invalid parameters, or API rate limits
- Polling loop times out - session may be stuck in a long-running state without progress
- Final answer retrieval fails - could be due to session cleanup or network issues after completion
## Source
Extracted from: [https://github.com/redhat-community-ai-tools/UnifAI.git](https://github.com/redhat-community-ai-tools/UnifAI.git)
Confidence: 0.95
@@ -0,0 +1,6 @@
# Commands: multi-agent-workflow-execution
## Available Commands
- `/skill multi-agent-workflow-execution` — Load this skill
- `/run multi-agent-workflow-execution` — Execute workflow
@@ -0,0 +1,10 @@
# Examples: multi-agent-workflow-execution
## Usage Example
```python
# How to use this skill
# Inputs: Blueprint ID or name (to identify the workflow to execute), User shortcut (authentication identifier for the user), User question or prompt (input to the workflow), Base URL of the UnifAI API (endpoint for session management), Polling interval (seconds between status checks during execution)
# Process: Resolve the blueprint ID from either direct ID or name lookup via the API, handling cases where the blueprint is not found or not unique → Create a new session from the resolved blueprint using the session creation endpoint → Submit the session with the user's prompt to start the multi-agent workflow execution
# Outputs: Final workflow result or answer (text or structured data), Session status (completed, failed, or cancelled), Error details if the workflow execution fails or times out
```
@@ -0,0 +1,35 @@
{
"name": "multi-agent-workflow-execution",
"version": "1.0.0",
"goal": "Execute multi-agent AI workflows defined in YAML blueprints by creating sessions, submitting user prompts, and polling for completion until final answers are returned.",
"inputs": [
"Blueprint ID or name (to identify the workflow to execute)",
"User shortcut (authentication identifier for the user)",
"User question or prompt (input to the workflow)",
"Base URL of the UnifAI API (endpoint for session management)",
"Polling interval (seconds between status checks during execution)"
],
"steps": [
"Resolve the blueprint ID from either direct ID or name lookup via the API, handling cases where the blueprint is not found or not unique",
"Create a new session from the resolved blueprint using the session creation endpoint",
"Submit the session with the user's prompt to start the multi-agent workflow execution",
"Poll the session status at regular intervals until the session completes, fails, or is cancelled",
"Retrieve and return the final answer from the completed workflow"
],
"outputs": [
"Final workflow result or answer (text or structured data)",
"Session status (completed, failed, or cancelled)",
"Error details if the workflow execution fails or times out"
],
"failure_modes": [
"Blueprint not found or not unique - script exits with an error listing available blueprints",
"Session creation fails - may be due to invalid blueprint ID, authentication issues, or rate limiting",
"Session submission fails - could be due to network issues, invalid parameters, or API rate limits",
"Polling loop times out - session may be stuck in a long-running state without progress",
"Final answer retrieval fails - could be due to session cleanup or network issues after completion"
],
"confidence": 0.95,
"explanation": "The UnifAI repository contains a concrete, reusable workflow pattern for executing multi-agent AI workflows. The scripts/execution_workflow.py script demonstrates a complete pipeline: resolving blueprints by ID or name, creating sessions from blueprints, submitting user prompts to start workflows, polling session status until completion, and retrieving final answers. This pattern can be adapted to any multi-agent workflow defined in the YAML blueprint system, making it reusable across different use cases and teams.",
"source_repo": "https://github.com/redhat-community-ai-tools/UnifAI.git",
"score": 1.0
}
@@ -0,0 +1,9 @@
# Tests: multi-agent-workflow-execution
## Test Checklist
- [ ] Workflow has at least 3 steps
- [ ] All inputs are defined
- [ ] All outputs are defined
- [ ] Failure modes are documented
- [ ] Skill can be loaded without errors
@@ -0,0 +1,94 @@
---
name: three-tier-evaluation-pipeline
version: 1.0.0
description: Run tasks through three evaluation tiers (Run, Trace, Thread) to produce
comprehensive reports with human-in-the-loop validation
inputs:
- query/input text for the task
- search results (for trace tier evaluation)
- evaluation criteria and thresholds
steps:
- 'Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph
engine) to generate initial outputs and results'
- 'Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined
criteria, generating detailed analysis and scoring'
- 'Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion,
approval, and iterative refinement of the output'
outputs:
- Final consolidated report combining results from all three tiers
- Detailed scores and metrics per tier
- Threaded discussion logs for human review and approval
tags: []
metadata:
source_repo: https://github.com/itszhaoziyan-n/AgentKit.git
extracted_at: ''
confidence: 0.95
---
# three-tier-evaluation-pipeline
Run tasks through three evaluation tiers (Run, Trace, Thread) to produce comprehensive reports with human-in-the-loop validation
## Setup
**Dependencies:**
```text
pip install langgraph>=0.3 langchain-core>=0.3 langchain-anthropic>=0.3 langfuse>=2.0 mcp[server]>=1.24 tenacity>=9.0 fastapi>=0.115 psycopg[binary]>=3.1
```
**Setup steps:**
1. Install dependencies with pip install -e .[dev]
1. Start infrastructure: docker compose up -d (PostgreSQL, Langfuse, MCP server)
1. Configure environment variables (DATABASE_URL, MCP_API_KEY, etc.)
1. Run the pipeline: python -m eval.runner --tiers run,thread,trace
## Key Files
- `eval/ - contains the three-tier evaluation logic`
- `scripts/ci_gate.py - threshold update and benchmark validation`
- `agentkit/runtime/ - LangGraph engine for state management and graph execution`
## Steps
1. Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results
2. Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring
3. Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output
## Implementation Details
```python
The eval/ directory implements Run, Trace, and Thread stages with configurable tiers
```
```python
Benchmark suite (40 test cases) validates the pipeline's reliability
```
```python
CI/CD workflows (ci.yml, eval-fast.yml, eval-trace.yml) orchestrate the evaluation pipeline
```
## Inputs
- query/input text for the task
- search results (for trace tier evaluation)
- evaluation criteria and thresholds
## Outputs
- Final consolidated report combining results from all three tiers
- Detailed scores and metrics per tier
- Threaded discussion logs for human review and approval
## Failure Modes
- If the Run tier fails (e.g., code execution error), the pipeline can retry but may produce incomplete outputs
- If the Trace tier LLM-as-Judge produces low-quality evaluations, the Thread tier may need additional human intervention
- Threshold mismatches between tiers could cause the pipeline to exit early or require manual adjustment
## Source
Extracted from: [https://github.com/itszhaoziyan-n/AgentKit.git](https://github.com/itszhaoziyan-n/AgentKit.git)
Confidence: 0.95
@@ -0,0 +1,6 @@
# Commands: three-tier-evaluation-pipeline
## Available Commands
- `/skill three-tier-evaluation-pipeline` — Load this skill
- `/run three-tier-evaluation-pipeline` — Execute workflow
@@ -0,0 +1,10 @@
# Examples: three-tier-evaluation-pipeline
## Usage Example
```python
# How to use this skill
# Inputs: query/input text for the task, search results (for trace tier evaluation), evaluation criteria and thresholds
# Process: Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results → Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring → Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output
# Outputs: Final consolidated report combining results from all three tiers, Detailed scores and metrics per tier, Threaded discussion logs for human review and approval
```
@@ -0,0 +1,29 @@
{
"name": "three-tier-evaluation-pipeline",
"version": "1.0.0",
"goal": "Run tasks through three evaluation tiers (Run, Trace, Thread) to produce comprehensive reports with human-in-the-loop validation",
"inputs": [
"query/input text for the task",
"search results (for trace tier evaluation)",
"evaluation criteria and thresholds"
],
"steps": [
"Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results",
"Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring",
"Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output"
],
"outputs": [
"Final consolidated report combining results from all three tiers",
"Detailed scores and metrics per tier",
"Threaded discussion logs for human review and approval"
],
"failure_modes": [
"If the Run tier fails (e.g., code execution error), the pipeline can retry but may produce incomplete outputs",
"If the Trace tier LLM-as-Judge produces low-quality evaluations, the Thread tier may need additional human intervention",
"Threshold mismatches between tiers could cause the pipeline to exit early or require manual adjustment"
],
"confidence": 0.95,
"explanation": "The AgentKit repository contains a production-ready three-tier evaluation pipeline (Run \u2192 Trace \u2192 Thread) that can be adapted to any task requiring multi-stage validation. This workflow uses LangGraph for orchestration and LangChain for tool integration, making it portable across different agent engineering scenarios. The pattern is reusable because it separates concerns into distinct stages with clear inputs/outputs, allowing teams to plug in different evaluation criteria or human reviewers as needed.",
"source_repo": "https://github.com/itszhaoziyan-n/AgentKit.git",
"score": 1.0
}
@@ -0,0 +1,9 @@
# Tests: three-tier-evaluation-pipeline
## Test Checklist
- [ ] Workflow has at least 3 steps
- [ ] All inputs are defined
- [ ] All outputs are defined
- [ ] Failure modes are documented
- [ ] Skill can be loaded without errors