ocas-fellowEmpirical experimentation engine. Invoked by Mentor to evaluate, compare, and promote improvements to OCAS skills, prompts, heuristics, and workflows using benchmark-driven experiments. Returns best variant result with lineage. Not user-invocable — called only by Mentor.
Install via ClawdBot CLI:
clawdbot install indigokarasu/ocas-fellowGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://github.com/<agent-handleAudited Oct 6, 2026 · audit v1.0
Generated Oct 6, 2026
After a prompt refactor, a customer-facing agent's task completion rate drops from 92% to 78%. Fellow is invoked by Mentor to run a controlled experiment comparing the current champion prompt against three challenger variants over a fixed benchmark. It establishes a fresh baseline, executes N trials per variant in isolated containers, and returns the winning variant with lineage so the promotion is reversible.
A support automation team wants to know whether a 'concise tone' heuristic or an 'empathetic tone' heuristic resolves tickets faster. Fellow runs each variant across a fixed set of ticket transcripts, applies deterministic and LLM-rubric graders, and reports trial pass rates plus token metrics. Mentor uses the evidence to promote the better heuristic without shipping both to production.
An R&D lab's research agent has two competing workflows for literature summarization. Fellow evaluates both against a standardized eval.yaml suite, measuring accuracy and token cost. The result gives the lab an objective, traceable recommendation instead of relying on intuition, while keeping the losing variant archived with full mutation lineage.
A platform team maintains dozens of coordinated agent skills (Mentor, Forge, Chronicle, Praxis). When a new skill variant is proposed, Fellow executes the benchmark and checks pass rates against promotion thresholds (≥0.85 with N≥5, or N≥15 for core execution skills). Only variants that clear the threshold are promoted, keeping the skill registry evidence-based.
A fintech firm needs to test a new compliance-checking heuristic without any risk to live accounts. Fellow runs challenger variants inside isolated ocas-inception containers with zero external side effects, then returns a CycleResult with per-grader breakdowns. Compliance officers review the lineage before any promotion reaches production.
Fellow is deployed as an internal-only engine behind Mentor, optimizing an organization's own agent skills and prompts. It is not sold directly; value is captured through higher agent reliability, lower token spend, and fewer regression incidents. Costs are absorbed as platform infrastructure while savings show up as improved OKRs.
A vendor offers Fellow-based experimentation as part of a managed AI agent service. Customers submit skill variants and benchmark suites; Fellow runs controlled experiments and returns promotion-ready results with lineage. The provider charges a subscription or per-experiment fee, positioning empirical evaluation as a differentiated quality guarantee.
Fellow is integrated into a marketplace where third-party developers publish agent skills. Before any skill can be promoted or sold, it must pass Fellow's benchmark thresholds, creating trust for buyers. The marketplace takes a revenue share on certified skills and charges developers for premium evaluation runs.
💬 Integration Tip
Fellow is not user-invocable, so integrate it strictly behind Mentor by wiring Mentor's experiment-requests directory to Fellow's CycleResult output path. Provide a real eval.yaml suite and compute budget up front, and ensure isolated container support (ocas-inception) so challenger runs have zero external side effects.
Scored Oct 6, 2026
Control desktop applications on Windows — launch, close, focus, resize, move windows, simulate keyboard/mouse input, manage processes, control VSCode, read clipboard, and capture screen info. Use when the user wants to interact with any running program, switch windows, type text, press shortcuts, open files in VSCode, manage running processes, or get system display information.
Conduct rigorous, adversarial code reviews with zero tolerance for mediocrity. Use when users ask to "critically review" my code or a PR, "critique my code", "find issues in my code", or "what's wrong with this code". Identifies security holes, lazy patterns, edge case failures, and bad practices across Python, R, JavaScript/TypeScript, SQL, and front-end code. Scrutinizes error handling, type safety, performance, accessibility, and code quality. Provides structured feedback with severity tiers (Blocking, Required, Suggestions) and specific, actionable recommendations.
Pragmatic coding standards for writing clean, maintainable code — naming, functions, structure, anti-patterns, and pre-edit safety checks. Use when writing new code, refactoring existing code, reviewing code quality, or establishing coding standards.
Claude Code integration for OpenClaw. This skill provides interfaces to: - Query Claude Code documentation from https://code.claude.com/docs - Manage subagents and coding tasks - Execute AI-assisted coding workflows - Access best practices and common workflows Use this skill when users want to: - Get help with coding tasks - Query Claude Code documentation - Manage AI-assisted development workflows - Execute complex programming tasks
Plan, draft, version, and refine written content with enforced versioning and quality audits.
Use when writing tests, creating test strategies, or building automation frameworks. Invoke for unit tests, integration tests, E2E, coverage analysis, performance testing, security testing.