custom-llm-provider-setupWire any OpenAI-compatible LLM API into Hermes Agent as a custom provider — Cloudflare Workers AI, self-hosted vLLM, Ollama, LM Studio, or any service that implements /v1/chat/completions.
Install via ClawdBot CLI:
clawdbot install mina-atef-00/custom-llm-provider-setupGrade Limited — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Sends data to undocumented external endpoint (potential exfiltration)
POST → https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/run/@cf/meta/llamaCalls external URL not in known-safe list
https://github.com/mina-atef-00/agent-skillsUses known external API (expected, informational)
api.github.comAudited Sep 21, 2026 · audit v1.0
Generated Sep 22, 2026
A healthcare or legal team runs Ollama or vLLM on internal GPU servers so no data leaves their network. They wire the endpoint into Hermes as a custom provider via `model.api_key` in config.yaml, then validate the model with a streaming chat completion probe before enabling the agent. This satisfies compliance requirements while still giving staff an agent loop.
A solo developer builds an internal agent using Cloudflare Workers AI's `@cf/` models, avoiding OpenRouter markup and paying only per-request. They verify each candidate model's callability with curl probes (handling 403 not-registered and 404 deployment-missing responses) before committing the model id and base URL into a named `providers:` entry.
An engineer prototyping agent workflows runs LM Studio on their laptop with a local model, pointing Hermes at `http://localhost:1234/v1` with a dummy API key. They iterate quickly on prompt and tool behavior with zero cloud spend, resetting the session after each config change, then swap to a hosted endpoint for production.
A platform team routes Hermes traffic through Portkey's OpenAI-compatible gateway to gain observability, rate limiting, and fallback across multiple upstream models. They configure a named provider entry once and switch `model.default` per request, using Hermes' context_length override to match each upstream model's window.
A research group needs a model with an unusually large context window (e.g. 262k for Kimi K2.6) that isn't in Hermes' built-in list. They set `model.context_length` explicitly and probe with streaming to confirm big-reasoning-model latency is acceptable before adding it to the agent's roster, avoiding silent truncation or timeouts.
A vendor deploys and operates OpenAI-compatible inference endpoints (vLLM/Ollama clusters) on behalf of clients, then hands over a base URL and API key pre-wired into Hermes provider config. Clients get data residency and predictable latency without running GPUs themselves.
A consultancy specializes in wiring bespoke or non-standard LLM endpoints into Hermes, including the host-gated credential workarounds, streaming probe validation, and multi-model roster decisions. They deliver a working config plus a probe report covering 403/404/200 outcomes per candidate model.
A startup resells access to a curated, live-probed roster of models through an OpenAI-compatible gateway compatible with Hermes. Customers pick a model or fallback chain via a `providers:` entry, and the platform guarantees each listed model has passed a `hermes chat -Q` smoke test.
💬 Integration Tip
Remember the host-gated credential guard — `OPENAI_API_KEY` is only forwarded to openai.com/azure hosts, so for Cloudflare/Ollama/vLLM endpoints you MUST put the key in `model.api_key` or a named `providers:` entry, and start a fresh Hermes session after any config change. Probe big reasoning models with `"stream": true` and validate usability with `hermes chat -Q --provider <p> -m <model> -q "..."`, since a bare non-streaming probe can time out misleadingly.
Scored Sep 22, 2026
Check Antigravity account quotas for Claude and Gemini models. Shows remaining quota and reset times with ban detection.
使用豆包(火山引擎)语音合成大模型 API 将文本转换为语音音频文件。支持声音复刻音色(S_ 开头的音色ID)和官方预置音色。当用户要求"语音合成"、"文字转语音"、"TTS"、"朗读文本"、"生成语音"、"用我的声音读"、"豆包语音"、"声音复刻合成"等相关请求时,务必使用此 skill。即使用户只是说"帮我把...
Intelligent model routing for sub-agent task delegation. Choose the optimal model based on task complexity, cost, and capability requirements. Reduces costs...
自动生成科技新闻摘要。从多个来源(RSS、Twitter、GitHub、Web Search)抓取科技新闻,整合后生成摘要。
让 AI 代理根据对话内容自动选择最合适的模型。四层识别(系统过滤→关键词→指示词→语义相似度),四池架构(高速/智能/人文/代理),五分支路由,全自动 Fallback 回路。支持 trigger_groups_all 非连续词组命中。
OpenAI API integration — chat completions, embeddings, image generation, audio transcription, file management, fine-tuning, and assistants via the OpenAI RES...