oh-my-knowledge evaluates AI knowledge artifacts
oh-my-knowledge is a local-first framework for comparing prompts, skills, RAG corpora, agent workflows, and runtime context while holding models and test samples constant. Why it matters: Prompt and retrieval changes can look better subjectively while reducing reliability. A repeatable evaluation harness helps decide whether an AI workflow artifact is ready to ship.
Try this: Run a small comparison between two prompt or retrieval variants, keep the test sample fixed, and inspect confidence intervals before adopting the higher-scoring version.