Files
Ty e5e179e541 Implement install stages S0–S2: bootstrap, env, model and vision smoke.
Operator can run make bootstrap/install through S2 using nemohermes/openshell wrappers; docs and implement queue updated. No S3+ and no push.
2026-07-27 11:47:34 -07:00

3.4 KiB

Hermes model configuration

Status: S2-ready — sufficient for operators configuring models.

Model slots

Slot Requirement Configured via
Main OpenAI-compatible chat + tools; endpoint may be outside Docker openshell inference set + .env
Auxiliary vision Required for media social; install smoke test .env + vision smoke
Other aux Default to same base/main endpoint unless overridden .env

Environment variables

All model configuration lives in .env. See .env.example for the template.

Variable Purpose Example
LUMINA_INFERENCE_BASE_URL OpenAI-compatible endpoint base URL http://192.168.1.203:8083/v1
LUMINA_INFERENCE_MODEL Main model identifier (must match what the server advertises) /home/ty/models/…/Qwen3.6-27B-UD-Q4_K_XL.gguf
LUMINA_INFERENCE_API_KEY API key for the inference endpoint (empty if unauthenticated) (leave empty for local endpoints)
LUMINA_VISION_MODEL Vision-capable model for social media / media understanding Same as main if multimodal
LUMINA_SANDBOX Sandbox name (used by nemohermes commands) hermes

Configuring the main model

The main model is configured through the OpenShell gateway:

openshell inference set \
  --provider compatible-endpoint \
  --model "$LUMINA_INFERENCE_MODEL" \
  --no-verify

This is done automatically by scripts/install/s2-models.sh when the gateway is connected.

Provider name: compatible-endpoint is the NemoClaw provider for OpenAI-compatible endpoints (llama.cpp, Ollama, vLLM, etc.).

Note: openshell inference set takes --provider and --model but not --url. The gateway resolves the endpoint URL from its own metadata (the compatible-endpoint provider reads the URL from the gateway config). The install script verifies the endpoint is reachable before calling openshell inference set.

Vision model

The vision model must support multimodal inputs (images). The install smoke test verifies this:

  1. Probes the models endpoint ($LUMINA_INFERENCE_BASE_URL/models) for a multimodal capability tag.
  2. Falls back to a chat completion probe if tags are not available.

If the main model is already multimodal, set LUMINA_VISION_MODEL to the same value as LUMINA_INFERENCE_MODEL.

Checking current configuration

# See what the gateway has configured
openshell inference get

# See what models the endpoint advertises
curl -s "$LUMINA_INFERENCE_BASE_URL/models" | jq '.models[].name'

# Check sandbox status
nemohermes <sandbox-name> status

Runtime changes

Per NemoClaw's runtime change matrix:

  • Inference route changes are typically hot (no rebuild needed).
  • Auxiliary model changes may require a gateway restart.
  • Always use nemohermes / openshell for mutations — never hand-edit in-sandbox config.

See design/DESIGN_PLAN.md §3.4 and NemoClaw runtime changes docs.

References