brainharness-autoresearchUse when user wants to optimize, improve, benchmark, or evaluate a skill's prompt. Triggers on "optimize skill", "improve skill prompt", "benchmark skill", "eval skill", "run autoresearch", "tune prompt", "prompt optimization", "skill evaluation", "A/B test prompt", "find best prompt", "auto-improve skill". Runs automated prompt experiments using the Karpathy autoresearch pattern.
Install via ClawdBot CLI:
clawdbot install zning1994/brainharness-autoresearchGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://github.com/zning1994/brainharness-autoresearchUses known external API (expected, informational)
api.anthropic.comAudited Oct 5, 2026 · audit v1.0
Generated Oct 4, 2026
An engineer responsible for an internal customer-support skill is frustrated by vague, apology-heavy replies. They build an eval.json with regex checks for ticket IDs and LLM-judged actionability, then let the autoresearch loop mutate the SKILL.md prompt over 30 experiments until pass rate climbs.
A platform team maintains several prompt variants for a code-review assistant and needs objective evidence for which performs best. They define word-count and contains evals, run autoresearch with --runs 5 for statistical significance, and compare scores in results.tsv before shipping the winner.
A search skill occasionally hallucinates placeholder URLs and returns thin summaries. The team adds banned_phrases and word_count evals to catch these failure modes, then uses autoresearch to find a prompt that keeps good behavior while eliminating the hallucination pattern.
A media-monitoring skill must handle both English and Chinese requests with equal quality. Using mixed-language test_inputs and a contains eval for localized summary markers, the team runs auto-improvement to discover a prompt that triggers consistent structured summaries in either language.
A skill author preparing to publish on the BrainHarness marketplace wants measurable proof of prompt quality. They build a small eval suite covering depth, sourcing, and tone, run autoresearch to derive an optimized prompt, and attach the results.tsv evidence to the listing.
The skill and eval tooling stay free and MIT-licensed to drive adoption, while a hosted service executes long autoresearch sweeps on managed GPU/LLM infrastructure and stores versioned results. Teams pay for compute hours and result history rather than for the code.
The skill is packaged into an internal platform that connects to a company's private skill registry, enforces eval standards, and runs scheduled optimization jobs with audit logs. VPC deployment and SSO justify a per-seat or per-workspace license.
Individual developers get a limited number of free experiments per month in a web UI that generates eval.json and visualizes dashboards. Power users and small teams upgrade to unlimited experiments, private eval suites, and API access billed by consumed tokens or experiment count.
💬 Integration Tip
Start with 3-5 eval checks and --runs 3 before scaling up, and always inspect the prompt_diff column so the winning mutation reflects a real improvement rather than an eval artifact.
Scored Oct 5, 2026
Medical device risk management specialist implementing ISO 14971 throughout product lifecycle. Provides risk analysis, risk evaluation, risk control, and pos...
When the user wants to plan a product launch, feature announcement, or release strategy. Also use when the user mentions 'launch,' 'Product Hunt,' 'feature r...
When the user wants to build a free tool for marketing — lead generation, SEO value, or brand awareness. Use when they mention 'engineering as marketing,' 'f...
管理多类型项目看板,支持新增项目、变更状态与版本、分类管理及查看项目总览和变更日志。
Competitor Analysis — SEO/GEO Intelligence & Market Positioning. Analyze competitor SEO rankings, AI search citations, content strategy, and market posi...
Conduct structured PHQ-9 depression symptom screening and submit the completed assessment for evaluation.