ab-test-runnerDesign and execute A/B testing experiments for LLM prompts, agent behaviors, and content production. Activate when user says "run an AB test", "design an exp...
Install via ClawdBot CLI:
clawdbot install demo112/ab-test-runnerGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated May 3, 2026
An e-commerce company wants to test whether a friendly vs. professional tone in AI chatbot responses improves customer satisfaction scores. They can use AB Test Runner to design an experiment with variant prompts, collect self and cross scores, and determine which style leads to higher ratings.
A marketing agency compares two versions of a product description: one with bullet points and one with narrative storytelling. The experiment measures engagement metrics like click-through rate and readability scores, helping the agency optimize content for different audiences.
A SaaS company tests whether an AI assistant that proactively suggests features vs. one that only responds to queries leads to higher user adoption. The experiment defines behavior prompts and evaluates user satisfaction and feature usage.
A content platform tests two text-to-speech speeds (normal vs. +15%) for audiobook narration to see which yields higher completion rates. The experiment uses A/B testing with listening samples and completion metrics.
Offer structured A/B testing for AI prompts, agent behaviors, and content as a paid service to companies optimizing their AI interactions. Revenue comes from per-experiment fees or monthly subscriptions.
Provide basic A/B testing for free (e.g., up to 5 experiments/month) and charge for advanced features like cross-scoring, detailed reports, and unlimited experiments. Revenue from premium tiers.
Offer consulting services to design and analyze experiments for clients, plus a tool that automates the process. Revenue from consulting fees and software licensing.
💬 Integration Tip
Integrate with your existing LLM API calls by wrapping them with the experiment manifest and result logging functions.
Scored Apr 19, 2026
Control desktop applications on Windows — launch, close, focus, resize, move windows, simulate keyboard/mouse input, manage processes, control VSCode, read clipboard, and capture screen info. Use when the user wants to interact with any running program, switch windows, type text, press shortcuts, open files in VSCode, manage running processes, or get system display information.
Conduct rigorous, adversarial code reviews with zero tolerance for mediocrity. Use when users ask to "critically review" my code or a PR, "critique my code", "find issues in my code", or "what's wrong with this code". Identifies security holes, lazy patterns, edge case failures, and bad practices across Python, R, JavaScript/TypeScript, SQL, and front-end code. Scrutinizes error handling, type safety, performance, accessibility, and code quality. Provides structured feedback with severity tiers (Blocking, Required, Suggestions) and specific, actionable recommendations.
Coding style memory that adapts to your preferences, conventions, and patterns for consistent coding.
Pragmatic coding standards for writing clean, maintainable code — naming, functions, structure, anti-patterns, and pre-edit safety checks. Use when writing new code, refactoring existing code, reviewing code quality, or establishing coding standards.
Claude Code integration for OpenClaw. This skill provides interfaces to: - Query Claude Code documentation from https://code.claude.com/docs - Manage subagents and coding tasks - Execute AI-assisted coding workflows - Access best practices and common workflows Use this skill when users want to: - Get help with coding tasks - Query Claude Code documentation - Manage AI-assisted development workflows - Execute complex programming tasks
Plan, draft, version, and refine written content with enforced versioning and quality audits.