Compare commits

..

1 Commits

Author SHA1 Message Date
Hermes Pipeline 1ba56fd7e3 Add Skill: three-tier-evaluation-pipeline
Extracted from: https://github.com/itszhaoziyan-n/AgentKit.git
Score: 1.0
2026-08-05 17:05:03 +00:00
9 changed files with 140 additions and 143 deletions
-96
View File
@@ -1,96 +0,0 @@
---
name: blacknode-graph-workflow
version: 1.0.0
description: Build and execute node-based AI workflows with LLM agents and processing
nodes
inputs:
- Model configuration (NIM_MODEL, API keys for NVIDIA NIM, OpenAI, Anthropic)
- Node definitions with specified inputs and outputs (Literal, LLMAgent, FileWrite,
etc.)
- Data sources (URLs, text content, or other inputs for the workflow)
steps:
- Initialize a blacknode.Graph instance to create the workflow structure
- Add nodes to the graph with defined inputs and outputs (e.g., Literal, LLMAgent,
FileWrite)
- Define edges connecting nodes to establish data flow between them
- Execute the graph using cook() to run the workflow and generate outputs
outputs:
- Processed results from the final node (e.g., printed text, written files, or generated
data)
- Graph execution status and any errors encountered during execution
tags: []
metadata:
source_repo: https://github.com/temiroff/Blacknode.git
extracted_at: ''
confidence: 0.95
---
# blacknode-graph-workflow
Build and execute node-based AI workflows with LLM agents and processing nodes
## Setup
**Dependencies:**
```text
pip install blacknode (core package) anthropic>=0.25 openai>=1.0 petgraph (for graph operations)
```
**Setup steps:**
1. Install blacknode package: pip install blacknode
1. Configure model API keys (NIM_API_KEY, OPENAI_API_KEY, etc.) in .env or editor
1. Create a Graph instance and add nodes with inputs/outputs
1. Define node connections in g._edges list
1. Execute with g.cook() to run the workflow and capture results
## Key Files
- `blacknode/blacknode.py (Graph class implementation)`
- `examples/hello_agent.py (simple LLM agent workflow)`
- `examples/converted_nvidia_nim.py (NIM model workflow)`
## Steps
1. Initialize a blacknode.Graph instance to create the workflow structure
2. Add nodes to the graph with defined inputs and outputs (e.g., Literal, LLMAgent, FileWrite)
3. Define edges connecting nodes to establish data flow between them
4. Execute the graph using cook() to run the workflow and generate outputs
## Implementation Details
```python
g = bn.Graph()
```
```python
g._edges = [{'from': 'model', 'from_port': 'value', 'to': 'agent', 'to_port': 'model'}]
```
```python
result = g.cook(output_node, 'value')
```
## Inputs
- Model configuration (NIM_MODEL, API keys for NVIDIA NIM, OpenAI, Anthropic)
- Node definitions with specified inputs and outputs (Literal, LLMAgent, FileWrite, etc.)
- Data sources (URLs, text content, or other inputs for the workflow)
## Outputs
- Processed results from the final node (e.g., printed text, written files, or generated data)
- Graph execution status and any errors encountered during execution
## Failure Modes
- Missing or invalid model API key causing graph initialization failure
- Incorrect node connections or missing edge definitions leading to runtime errors
- Model not found or unavailable in the specified environment causing execution failure
- Graph edges not properly defined or mismatched causing cook() to fail
## Source
Extracted from: [https://github.com/temiroff/Blacknode.git](https://github.com/temiroff/Blacknode.git)
Confidence: 0.95
@@ -1,6 +0,0 @@
# Commands: blacknode-graph-workflow
## Available Commands
- `/skill blacknode-graph-workflow` — Load this skill
- `/run blacknode-graph-workflow` — Execute workflow
@@ -1,10 +0,0 @@
# Examples: blacknode-graph-workflow
## Usage Example
```python
# How to use this skill
# Inputs: Model configuration (NIM_MODEL, API keys for NVIDIA NIM, OpenAI, Anthropic), Node definitions with specified inputs and outputs (Literal, LLMAgent, FileWrite, etc.), Data sources (URLs, text content, or other inputs for the workflow)
# Process: Initialize a blacknode.Graph instance to create the workflow structure → Add nodes to the graph with defined inputs and outputs (e.g., Literal, LLMAgent, FileWrite) → Define edges connecting nodes to establish data flow between them
# Outputs: Processed results from the final node (e.g., printed text, written files, or generated data), Graph execution status and any errors encountered during execution
```
@@ -1,30 +0,0 @@
{
"name": "blacknode-graph-workflow",
"version": "1.0.0",
"goal": "Build and execute node-based AI workflows with LLM agents and processing nodes",
"inputs": [
"Model configuration (NIM_MODEL, API keys for NVIDIA NIM, OpenAI, Anthropic)",
"Node definitions with specified inputs and outputs (Literal, LLMAgent, FileWrite, etc.)",
"Data sources (URLs, text content, or other inputs for the workflow)"
],
"steps": [
"Initialize a blacknode.Graph instance to create the workflow structure",
"Add nodes to the graph with defined inputs and outputs (e.g., Literal, LLMAgent, FileWrite)",
"Define edges connecting nodes to establish data flow between them",
"Execute the graph using cook() to run the workflow and generate outputs"
],
"outputs": [
"Processed results from the final node (e.g., printed text, written files, or generated data)",
"Graph execution status and any errors encountered during execution"
],
"failure_modes": [
"Missing or invalid model API key causing graph initialization failure",
"Incorrect node connections or missing edge definitions leading to runtime errors",
"Model not found or unavailable in the specified environment causing execution failure",
"Graph edges not properly defined or mismatched causing cook() to fail"
],
"confidence": 0.95,
"explanation": "Blacknode provides a standardized Graph-based workflow pattern where users create node graphs using the blacknode.Graph class. This pattern is reusable across projects as it follows a consistent structure: initialize a graph, add nodes with defined inputs/outputs, connect them with edges, and execute with cook(). The examples demonstrate this pattern with LLM agents and text processing pipelines, making it adaptable to various robotics and AI workflows.",
"source_repo": "https://github.com/temiroff/Blacknode.git",
"score": 1.0
}
@@ -0,0 +1,94 @@
---
name: three-tier-evaluation-pipeline
version: 1.0.0
description: Run tasks through three evaluation tiers (Run, Trace, Thread) to produce
comprehensive reports with human-in-the-loop validation
inputs:
- query/input text for the task
- search results (for trace tier evaluation)
- evaluation criteria and thresholds
steps:
- 'Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph
engine) to generate initial outputs and results'
- 'Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined
criteria, generating detailed analysis and scoring'
- 'Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion,
approval, and iterative refinement of the output'
outputs:
- Final consolidated report combining results from all three tiers
- Detailed scores and metrics per tier
- Threaded discussion logs for human review and approval
tags: []
metadata:
source_repo: https://github.com/itszhaoziyan-n/AgentKit.git
extracted_at: ''
confidence: 0.95
---
# three-tier-evaluation-pipeline
Run tasks through three evaluation tiers (Run, Trace, Thread) to produce comprehensive reports with human-in-the-loop validation
## Setup
**Dependencies:**
```text
pip install langgraph>=0.3 langchain-core>=0.3 langchain-anthropic>=0.3 langfuse>=2.0 mcp[server]>=1.24 tenacity>=9.0 fastapi>=0.115 psycopg[binary]>=3.1
```
**Setup steps:**
1. Install dependencies with pip install -e .[dev]
1. Start infrastructure: docker compose up -d (PostgreSQL, Langfuse, MCP server)
1. Configure environment variables (DATABASE_URL, MCP_API_KEY, etc.)
1. Run the pipeline: python -m eval.runner --tiers run,thread,trace
## Key Files
- `eval/ - contains the three-tier evaluation logic`
- `scripts/ci_gate.py - threshold update and benchmark validation`
- `agentkit/runtime/ - LangGraph engine for state management and graph execution`
## Steps
1. Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results
2. Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring
3. Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output
## Implementation Details
```python
The eval/ directory implements Run, Trace, and Thread stages with configurable tiers
```
```python
Benchmark suite (40 test cases) validates the pipeline's reliability
```
```python
CI/CD workflows (ci.yml, eval-fast.yml, eval-trace.yml) orchestrate the evaluation pipeline
```
## Inputs
- query/input text for the task
- search results (for trace tier evaluation)
- evaluation criteria and thresholds
## Outputs
- Final consolidated report combining results from all three tiers
- Detailed scores and metrics per tier
- Threaded discussion logs for human review and approval
## Failure Modes
- If the Run tier fails (e.g., code execution error), the pipeline can retry but may produce incomplete outputs
- If the Trace tier LLM-as-Judge produces low-quality evaluations, the Thread tier may need additional human intervention
- Threshold mismatches between tiers could cause the pipeline to exit early or require manual adjustment
## Source
Extracted from: [https://github.com/itszhaoziyan-n/AgentKit.git](https://github.com/itszhaoziyan-n/AgentKit.git)
Confidence: 0.95
@@ -0,0 +1,6 @@
# Commands: three-tier-evaluation-pipeline
## Available Commands
- `/skill three-tier-evaluation-pipeline` — Load this skill
- `/run three-tier-evaluation-pipeline` — Execute workflow
@@ -0,0 +1,10 @@
# Examples: three-tier-evaluation-pipeline
## Usage Example
```python
# How to use this skill
# Inputs: query/input text for the task, search results (for trace tier evaluation), evaluation criteria and thresholds
# Process: Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results → Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring → Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output
# Outputs: Final consolidated report combining results from all three tiers, Detailed scores and metrics per tier, Threaded discussion logs for human review and approval
```
@@ -0,0 +1,29 @@
{
"name": "three-tier-evaluation-pipeline",
"version": "1.0.0",
"goal": "Run tasks through three evaluation tiers (Run, Trace, Thread) to produce comprehensive reports with human-in-the-loop validation",
"inputs": [
"query/input text for the task",
"search results (for trace tier evaluation)",
"evaluation criteria and thresholds"
],
"steps": [
"Step 1: Execute the main task using the Run tier of the evaluation pipeline (agentkit/runtime/LangGraph engine) to generate initial outputs and results",
"Step 2: Run the Trace tier where an LLM-as-Judge evaluates the output against defined criteria, generating detailed analysis and scoring",
"Step 3: Execute the Thread tier which facilitates human-in-the-loop discussion, approval, and iterative refinement of the output"
],
"outputs": [
"Final consolidated report combining results from all three tiers",
"Detailed scores and metrics per tier",
"Threaded discussion logs for human review and approval"
],
"failure_modes": [
"If the Run tier fails (e.g., code execution error), the pipeline can retry but may produce incomplete outputs",
"If the Trace tier LLM-as-Judge produces low-quality evaluations, the Thread tier may need additional human intervention",
"Threshold mismatches between tiers could cause the pipeline to exit early or require manual adjustment"
],
"confidence": 0.95,
"explanation": "The AgentKit repository contains a production-ready three-tier evaluation pipeline (Run \u2192 Trace \u2192 Thread) that can be adapted to any task requiring multi-stage validation. This workflow uses LangGraph for orchestration and LangChain for tool integration, making it portable across different agent engineering scenarios. The pattern is reusable because it separates concerns into distinct stages with clear inputs/outputs, allowing teams to plug in different evaluation criteria or human reviewers as needed.",
"source_repo": "https://github.com/itszhaoziyan-n/AgentKit.git",
"score": 1.0
}
@@ -1,4 +1,4 @@
# Tests: blacknode-graph-workflow
# Tests: three-tier-evaluation-pipeline
## Test Checklist