Open-source software for human-in-the-loop AI.
HITL Kit
Human-in-the-loop AI, measured properly.
v0.6 · deployed at hitlkit.dev · 16 primitives via shadcn CLI · six packages on npm.
eval-kit
A measurement instrument for multi-step research agents.
v0.3.1 stable · @eval-kit/core, @eval-kit/ui, @eval-kit/seed-suite on npm · three reference suites (research, coding, support) · four adapters (anthropic, openai, http, mock) · file-based, single-user, not a hosted service.
tag-kit
Structured tagging primitives for human-in-the-loop annotation workflows.
@tag-kit/core (zero runtime deps) · @tag-kit/ui (headless React) · extracted from a real moderation app (inertial).
Collapse
A Claude Code skill-building framework.
v0.2 · active development. Shipped: MDX ingestor with 21 cross-stack reference lessons, notebook ingestor with admonition auto-prefill, template engine with cross-language equivalents + trigger-phrase derivation, atomic persistence, three-tier skill quality linter, and MCP server scaffold generation — one annotation can now become a skill or a working MCP tool.
Hologram
Live observability, guided skills, and an agent (MCP) surface for Blender → glTF pipelines.
v0.6.0 · on PyPI as hologram-gltf · Python 3.10+ · MIT.
Human-in-the-loop AI, measured properly.
Evaluate AI on whether it assists humans without displacing them — not on whether it can finish the task alone.
Read the paper: An AI Measurement ProblemThe gate is the unit of measurement
№ 005·published
A reframing of what the akaOSS measurement family actually measures: not task completion, but the gate, the moment an agent hands control back to a human. Gates come in two kinds that must never share a metric, autonomous benchmarks structurally penalize both, and a gate over unverifiable output is oversight theater. Regulation now mandates effective human oversight; nothing measures whether oversight is effective. That gap is the research program.
All findings