zouroboros-benchBenchmark harness for AI memory systems. Evaluates LongMemEval, LoCoMo, and ConvoMem datasets against any memory backend via the zouroboros-memory CLI. Inclu...
Install via ClawdBot CLI:
clawdbot install marlandoj/zouroboros-benchGrade Limited — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Accesses sensitive credential files or environment variables
process.env.OPENAIPotentially destructive shell commands in tool definitions
eval (Calls external URL not in known-safe list
https://github.com/AlaricHQ/zouroboros-openclawUses known external API (expected, informational)
api.openai.comGenerated May 9, 2026
AI agents using persistent memory (e.g., customer support bots) must be tested for recall accuracy over long conversations. Zouroboros Bench runs LongMemEval, LoCoMo, and ConvoMem to detect memory drift before deployment.
Financial or healthcare AI must retain conversation context without error. Benchmarks track memory consistency across sessions, flagging deviations that could lead to compliance failures.
Organizations evaluating different memory backends (e.g., SQLite vs. hosted vector stores) can use Zouroboros Bench with any CLI-compatible memory system to compare accuracy and latency.
As AI agents evolve, their memory subsystems may degrade. The Mimir judge continuously evaluates runs against baseline performance, alerting teams to regressions.
Companies using offline LLMs (via Ollama) can benchmark memory without exposing data to external APIs, ensuring privacy while maintaining performance metrics.
Free tool adoption drives ecosystem growth; commercial licenses or priority support for enterprise customers generate revenue.
Offer managed benchmarking runs with dashboards and historical trend analysis on a subscription basis.
Tailor Zouroboros Bench for specific client memory backends or proprietary datasets, delivered as professional services.
💬 Integration Tip
Simply install via npm and set OPENAI_API_KEY for the judge; use the provided CLI commands to start benchmarking immediately.
Scored May 9, 2026
Audited Apr 17, 2026 · audit v1.0
Structured reasoning through sequential thinking — break complex problems into steps, solve each independently, verify consistency, synthesize conclusions wi...
Loads and manages company context for all C-suite advisor skills. Reads ~/.claude/company-context.md, detects stale context (>90 days), enriches context duri...
Automatically prune and compact agent memory files to prevent unbounded growth. Circular buffer for logs, importance-based retention for state, and configura...
Founder onboarding interview that captures company context across 7 dimensions. Invoke with /cs:setup for initial interview or /cs:update for quarterly refre...
Ultimate AI agent memory system for Cursor, Claude, ChatGPT & Copilot. WAL protocol + vector search + git-notes + cloud backup. Never lose context again. Vib...
Expert skill for memory-lancedb-pro — a production-grade LanceDB-backed long-term memory plugin for OpenClaw agents with hybrid retrieval, cross-encoder rera...