Compare commits
1 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| e20b5370e1 |
@@ -0,0 +1,77 @@
|
||||
---
|
||||
name: literature-review-with-traceable-ai-evidence
|
||||
version: 1.0.0
|
||||
description: Enable researchers to ingest documents, asynchronously index them, and
|
||||
interact with multi-agent AI to answer questions with verifiable citations to original
|
||||
text.
|
||||
inputs:
|
||||
- Documents in PDF, Office, image, or text formats
|
||||
- Research questions or topics of interest
|
||||
- 'Optional: user model configuration via .env or settings'
|
||||
steps:
|
||||
- Upload documents to a project (via desktop app or web UI)
|
||||
- System asynchronously converts Office docs to PDF if needed, runs OCR to extract
|
||||
text with coordinates, chunks and embeds into vector store
|
||||
- User starts a main research session or creates exploration branches without waiting
|
||||
for indexing to finish
|
||||
- Leader agent receives query and delegates subtasks to researcher, reviewer, writer
|
||||
subagents
|
||||
- Subagents perform hybrid retrieval and rerank to find relevant chunks with source
|
||||
coordinates
|
||||
- Agents synthesize answers and return evidence with clickable citations that highlight
|
||||
original pages
|
||||
- User verifies conclusions by navigating to cited source locations and can save notes
|
||||
to research memory or mind map
|
||||
outputs:
|
||||
- AI-generated answers with traceable evidence (coordinates, page highlights)
|
||||
- Research session history with branches
|
||||
- Indexed document library for future queries
|
||||
- Mind maps or structured notes
|
||||
- Persistent run events for resuming sessions
|
||||
tags: []
|
||||
metadata:
|
||||
source_repo: https://github.com/0verL1nk/PaperSage.git
|
||||
extracted_at: ''
|
||||
confidence: 0.85
|
||||
---
|
||||
|
||||
# literature-review-with-traceable-ai-evidence
|
||||
|
||||
Enable researchers to ingest documents, asynchronously index them, and interact with multi-agent AI to answer questions with verifiable citations to original text.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Upload documents to a project (via desktop app or web UI)
|
||||
2. System asynchronously converts Office docs to PDF if needed, runs OCR to extract text with coordinates, chunks and embeds into vector store
|
||||
3. User starts a main research session or creates exploration branches without waiting for indexing to finish
|
||||
4. Leader agent receives query and delegates subtasks to researcher, reviewer, writer subagents
|
||||
5. Subagents perform hybrid retrieval and rerank to find relevant chunks with source coordinates
|
||||
6. Agents synthesize answers and return evidence with clickable citations that highlight original pages
|
||||
7. User verifies conclusions by navigating to cited source locations and can save notes to research memory or mind map
|
||||
|
||||
## Inputs
|
||||
|
||||
- Documents in PDF, Office, image, or text formats
|
||||
- Research questions or topics of interest
|
||||
- Optional: user model configuration via .env or settings
|
||||
|
||||
## Outputs
|
||||
|
||||
- AI-generated answers with traceable evidence (coordinates, page highlights)
|
||||
- Research session history with branches
|
||||
- Indexed document library for future queries
|
||||
- Mind maps or structured notes
|
||||
- Persistent run events for resuming sessions
|
||||
|
||||
## Failure Modes
|
||||
|
||||
- Missing local Office/LibreOffice converter causes document conversion failure
|
||||
- First-time model download may be slow or require network
|
||||
- OCR may have low confidence on poor quality scans
|
||||
- Retrieval might miss context if chunking splits semantics
|
||||
- Multi-agent coordination could produce conflicting intermediate results
|
||||
|
||||
## Source
|
||||
|
||||
Extracted from: [https://github.com/0verL1nk/PaperSage.git](https://github.com/0verL1nk/PaperSage.git)
|
||||
Confidence: 0.85
|
||||
@@ -0,0 +1,6 @@
|
||||
# Commands: literature-review-with-traceable-ai-evidence
|
||||
|
||||
## Available Commands
|
||||
|
||||
- `/skill literature-review-with-traceable-ai-evidence` — Load this skill
|
||||
- `/run literature-review-with-traceable-ai-evidence` — Execute workflow
|
||||
@@ -0,0 +1,10 @@
|
||||
# Examples: literature-review-with-traceable-ai-evidence
|
||||
|
||||
## Usage Example
|
||||
|
||||
```python
|
||||
# How to use this skill
|
||||
# Inputs: Documents in PDF, Office, image, or text formats, Research questions or topics of interest, Optional: user model configuration via .env or settings
|
||||
# Process: Upload documents to a project (via desktop app or web UI) → System asynchronously converts Office docs to PDF if needed, runs OCR to extract text with coordinates, chunks and embeds into vector store → User starts a main research session or creates exploration branches without waiting for indexing to finish
|
||||
# Outputs: AI-generated answers with traceable evidence (coordinates, page highlights), Research session history with branches, Indexed document library for future queries, Mind maps or structured notes, Persistent run events for resuming sessions
|
||||
```
|
||||
@@ -0,0 +1,37 @@
|
||||
{
|
||||
"name": "literature-review-with-traceable-ai-evidence",
|
||||
"version": "1.0.0",
|
||||
"goal": "Enable researchers to ingest documents, asynchronously index them, and interact with multi-agent AI to answer questions with verifiable citations to original text.",
|
||||
"inputs": [
|
||||
"Documents in PDF, Office, image, or text formats",
|
||||
"Research questions or topics of interest",
|
||||
"Optional: user model configuration via .env or settings"
|
||||
],
|
||||
"steps": [
|
||||
"Upload documents to a project (via desktop app or web UI)",
|
||||
"System asynchronously converts Office docs to PDF if needed, runs OCR to extract text with coordinates, chunks and embeds into vector store",
|
||||
"User starts a main research session or creates exploration branches without waiting for indexing to finish",
|
||||
"Leader agent receives query and delegates subtasks to researcher, reviewer, writer subagents",
|
||||
"Subagents perform hybrid retrieval and rerank to find relevant chunks with source coordinates",
|
||||
"Agents synthesize answers and return evidence with clickable citations that highlight original pages",
|
||||
"User verifies conclusions by navigating to cited source locations and can save notes to research memory or mind map"
|
||||
],
|
||||
"outputs": [
|
||||
"AI-generated answers with traceable evidence (coordinates, page highlights)",
|
||||
"Research session history with branches",
|
||||
"Indexed document library for future queries",
|
||||
"Mind maps or structured notes",
|
||||
"Persistent run events for resuming sessions"
|
||||
],
|
||||
"failure_modes": [
|
||||
"Missing local Office/LibreOffice converter causes document conversion failure",
|
||||
"First-time model download may be slow or require network",
|
||||
"OCR may have low confidence on poor quality scans",
|
||||
"Retrieval might miss context if chunking splits semantics",
|
||||
"Multi-agent coordination could produce conflicting intermediate results"
|
||||
],
|
||||
"confidence": 0.85,
|
||||
"explanation": "The README describes PaperSage's core workflow: asynchronous document ingestion with OCR/indexing, followed by multi-agent question answering with cited evidence. This process is not tied to the specific codebase and can be reused as a general literature review methodology for any document-centric research using RAG and agent collaboration.",
|
||||
"source_repo": "https://github.com/0verL1nk/PaperSage.git",
|
||||
"score": 1.0
|
||||
}
|
||||
+1
-1
@@ -1,4 +1,4 @@
|
||||
# Tests: local-document-research-with-traceable-citations
|
||||
# Tests: literature-review-with-traceable-ai-evidence
|
||||
|
||||
## Test Checklist
|
||||
|
||||
@@ -1,82 +0,0 @@
|
||||
---
|
||||
name: local-document-research-with-traceable-citations
|
||||
version: 1.0.0
|
||||
description: Enable users to import local documents, asynchronously process them into
|
||||
an indexed knowledge base, and obtain AI-generated answers that cite specific page
|
||||
locations and OCR evidence.
|
||||
inputs:
|
||||
- Local document files (PDF, DOCX, PPTX, XLSX, images, TXT)
|
||||
- Configured LLM API endpoint and keys (via .env or settings)
|
||||
- Optional web search service config if enabled
|
||||
- Local OCR model cache (downloaded on first use)
|
||||
steps:
|
||||
- Import documents into a project; files are queued for asynchronous processing.
|
||||
- Convert non-PDF formats (DOCX, PPTX, XLSX) to PDF using native Office or LibreOffice
|
||||
fallback.
|
||||
- Run OCR (PaddleOCR) on PDF pages/images to extract text, page numbers, polygons,
|
||||
and confidence scores.
|
||||
- Chunk text and generate embeddings; publish to LanceDB hybrid index (dense vector
|
||||
+ full-text) only when fully processed.
|
||||
- User starts a research session or branch; Leader agent analyzes query.
|
||||
- Hybrid RAG retrieves candidate chunks from ready documents; dynamic material scope
|
||||
ensures no half-indexed docs.
|
||||
- Leader delegates tasks to sub-agents (researcher, reviewer, writer) via constrained
|
||||
task capability; each delegation logs start, completion, duration, and evidence.
|
||||
- Generate answer that references only actually used evidence; citations include document
|
||||
ID, page, coordinates.
|
||||
- User clicks citation to open original document and view highlighted OCR location.
|
||||
outputs:
|
||||
- Project with indexed document library (SQLite metadata + LanceDB vectors)
|
||||
- AI answers with verifiable citations to source pages
|
||||
- Evidence preview with page image and OCR highlight polygons
|
||||
- Persistent session history, branches, and long-term memory
|
||||
tags: []
|
||||
metadata:
|
||||
source_repo: https://github.com/0verL1nk/PaperSage.git
|
||||
extracted_at: ''
|
||||
confidence: 0.85
|
||||
---
|
||||
|
||||
# local-document-research-with-traceable-citations
|
||||
|
||||
Enable users to import local documents, asynchronously process them into an indexed knowledge base, and obtain AI-generated answers that cite specific page locations and OCR evidence.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Import documents into a project; files are queued for asynchronous processing.
|
||||
2. Convert non-PDF formats (DOCX, PPTX, XLSX) to PDF using native Office or LibreOffice fallback.
|
||||
3. Run OCR (PaddleOCR) on PDF pages/images to extract text, page numbers, polygons, and confidence scores.
|
||||
4. Chunk text and generate embeddings; publish to LanceDB hybrid index (dense vector + full-text) only when fully processed.
|
||||
5. User starts a research session or branch; Leader agent analyzes query.
|
||||
6. Hybrid RAG retrieves candidate chunks from ready documents; dynamic material scope ensures no half-indexed docs.
|
||||
7. Leader delegates tasks to sub-agents (researcher, reviewer, writer) via constrained task capability; each delegation logs start, completion, duration, and evidence.
|
||||
8. Generate answer that references only actually used evidence; citations include document ID, page, coordinates.
|
||||
9. User clicks citation to open original document and view highlighted OCR location.
|
||||
|
||||
## Inputs
|
||||
|
||||
- Local document files (PDF, DOCX, PPTX, XLSX, images, TXT)
|
||||
- Configured LLM API endpoint and keys (via .env or settings)
|
||||
- Optional web search service config if enabled
|
||||
- Local OCR model cache (downloaded on first use)
|
||||
|
||||
## Outputs
|
||||
|
||||
- Project with indexed document library (SQLite metadata + LanceDB vectors)
|
||||
- AI answers with verifiable citations to source pages
|
||||
- Evidence preview with page image and OCR highlight polygons
|
||||
- Persistent session history, branches, and long-term memory
|
||||
|
||||
## Failure Modes
|
||||
|
||||
- OCR quality low for scanned images leading to poor extraction
|
||||
- Missing Office/LibreOffice causes conversion failure for Office docs
|
||||
- Interrupted indexing leaves documents unpublished and excluded from retrieval
|
||||
- LLM API outage or misconfiguration yields no answer
|
||||
- Citation coordinates mismatch due to chunk drift
|
||||
- Sub-agent recursion if constraints not enforced
|
||||
|
||||
## Source
|
||||
|
||||
Extracted from: [https://github.com/0verL1nk/PaperSage.git](https://github.com/0verL1nk/PaperSage.git)
|
||||
Confidence: 0.85
|
||||
@@ -1,6 +0,0 @@
|
||||
# Commands: local-document-research-with-traceable-citations
|
||||
|
||||
## Available Commands
|
||||
|
||||
- `/skill local-document-research-with-traceable-citations` — Load this skill
|
||||
- `/run local-document-research-with-traceable-citations` — Execute workflow
|
||||
@@ -1,10 +0,0 @@
|
||||
# Examples: local-document-research-with-traceable-citations
|
||||
|
||||
## Usage Example
|
||||
|
||||
```python
|
||||
# How to use this skill
|
||||
# Inputs: Local document files (PDF, DOCX, PPTX, XLSX, images, TXT), Configured LLM API endpoint and keys (via .env or settings), Optional web search service config if enabled, Local OCR model cache (downloaded on first use)
|
||||
# Process: Import documents into a project; files are queued for asynchronous processing. → Convert non-PDF formats (DOCX, PPTX, XLSX) to PDF using native Office or LibreOffice fallback. → Run OCR (PaddleOCR) on PDF pages/images to extract text, page numbers, polygons, and confidence scores.
|
||||
# Outputs: Project with indexed document library (SQLite metadata + LanceDB vectors), AI answers with verifiable citations to source pages, Evidence preview with page image and OCR highlight polygons, Persistent session history, branches, and long-term memory
|
||||
```
|
||||
@@ -1,40 +0,0 @@
|
||||
{
|
||||
"name": "local-document-research-with-traceable-citations",
|
||||
"version": "1.0.0",
|
||||
"goal": "Enable users to import local documents, asynchronously process them into an indexed knowledge base, and obtain AI-generated answers that cite specific page locations and OCR evidence.",
|
||||
"inputs": [
|
||||
"Local document files (PDF, DOCX, PPTX, XLSX, images, TXT)",
|
||||
"Configured LLM API endpoint and keys (via .env or settings)",
|
||||
"Optional web search service config if enabled",
|
||||
"Local OCR model cache (downloaded on first use)"
|
||||
],
|
||||
"steps": [
|
||||
"Import documents into a project; files are queued for asynchronous processing.",
|
||||
"Convert non-PDF formats (DOCX, PPTX, XLSX) to PDF using native Office or LibreOffice fallback.",
|
||||
"Run OCR (PaddleOCR) on PDF pages/images to extract text, page numbers, polygons, and confidence scores.",
|
||||
"Chunk text and generate embeddings; publish to LanceDB hybrid index (dense vector + full-text) only when fully processed.",
|
||||
"User starts a research session or branch; Leader agent analyzes query.",
|
||||
"Hybrid RAG retrieves candidate chunks from ready documents; dynamic material scope ensures no half-indexed docs.",
|
||||
"Leader delegates tasks to sub-agents (researcher, reviewer, writer) via constrained task capability; each delegation logs start, completion, duration, and evidence.",
|
||||
"Generate answer that references only actually used evidence; citations include document ID, page, coordinates.",
|
||||
"User clicks citation to open original document and view highlighted OCR location."
|
||||
],
|
||||
"outputs": [
|
||||
"Project with indexed document library (SQLite metadata + LanceDB vectors)",
|
||||
"AI answers with verifiable citations to source pages",
|
||||
"Evidence preview with page image and OCR highlight polygons",
|
||||
"Persistent session history, branches, and long-term memory"
|
||||
],
|
||||
"failure_modes": [
|
||||
"OCR quality low for scanned images leading to poor extraction",
|
||||
"Missing Office/LibreOffice causes conversion failure for Office docs",
|
||||
"Interrupted indexing leaves documents unpublished and excluded from retrieval",
|
||||
"LLM API outage or misconfiguration yields no answer",
|
||||
"Citation coordinates mismatch due to chunk drift",
|
||||
"Sub-agent recursion if constraints not enforced"
|
||||
],
|
||||
"confidence": 0.85,
|
||||
"explanation": "The README outlines a clear pipeline from document import to cited answer with evidence location, which is a reusable pattern for local-first RAG applications requiring traceability.",
|
||||
"source_repo": "https://github.com/0verL1nk/PaperSage.git",
|
||||
"score": 1.0
|
||||
}
|
||||
Reference in New Issue
Block a user