ah-llm-architectExpert LLM architect specializing in large language model architecture, deployment, and optimization. Masters LLM system design, fine-tuning strategies, and...
Install via ClawdBot CLI:
clawdbot install mtsatryan/ah-llm-architectGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated May 19, 2026
Deploy a scalable LLM-powered chatbot for customer support with low latency and high accuracy. The system integrates RAG for knowledge base retrieval and uses safety filters to ensure compliant responses.
Build a document summarization service for legal and medical firms using fine-tuned LLMs. Optimization focuses on long-context handling and cost-efficient inference via 4-bit quantization.
Implement a multi-model LLM system to detect toxic content, hate speech, and misinformation in user-generated content. Uses speculative decoding for sub-200ms latency and score-based routing to specialist models.
Create an adaptive tutoring platform that uses instruction-tuned LLMs with RLHF to provide personalized explanations and exercises. Incorporates retrieval augmentation to pull from curriculum materials.
Deploy a code review assistant using a fine-tuned CodeLLaMA model with chain-of-thought prompting. Integrates with CI/CD pipelines and uses token optimization to reduce API costs.
Charge enterprises a monthly fee per user or API key for access to custom LLM endpoints. Bundles include usage tiers with latency and throughput SLAs.
Offer a hosted LLM API with consumption-based pricing per token or per request. Provide fine-tuning and RAG services as add-ons.
Provide end-to-end model customization, including dataset preparation, LoRA training, and deployment. Charge a project fee plus ongoing maintenance retainer.
💬 Integration Tip
Use the communication protocol to exchange requirements with other agents; start with a simple serving setup (vLLM) and optimize iteratively based on monitoring.
Scored May 19, 2026
Use CodexBar CLI local cost usage to summarize per-model usage for Codex or Claude, including the current (most recent) model or a full model breakdown. Trigger when asked for model-level usage/cost data from codexbar, or when you need a scriptable per-model summary from codexbar cost JSON.
Check Antigravity account quotas for Claude and Gemini models. Shows remaining quota and reset times with ban detection.
使用豆包(火山引擎)语音合成大模型 API 将文本转换为语音音频文件。支持声音复刻音色(S_ 开头的音色ID)和官方预置音色。当用户要求"语音合成"、"文字转语音"、"TTS"、"朗读文本"、"生成语音"、"用我的声音读"、"豆包语音"、"声音复刻合成"等相关请求时,务必使用此 skill。即使用户只是说"帮我把...
Intelligent model routing for sub-agent task delegation. Choose the optimal model based on task complexity, cost, and capability requirements. Reduces costs...
自动生成科技新闻摘要。从多个来源(RSS、Twitter、GitHub、Web Search)抓取科技新闻,整合后生成摘要。
让 AI 代理根据对话内容自动选择最合适的模型。四层识别(系统过滤→关键词→指示词→语义相似度),四池架构(高速/智能/人文/代理),五分支路由,全自动 Fallback 回路。支持 trigger_groups_all 非连续词组命中。