Open-source tools for human-in-the-loop AI.
Projects
all projectsHITL Kit
Twenty React primitives for the moment an agent hands control to a person: interrupts, approvals, evidence, and provenance.
MeasurementDeployed at hitlkit.dev · every primitive installs individually via the shadcn CLI · copy, paste, own.
eval-kit
Scores whether your agent respects human authority: stops when it must, asks when it should, never averaged into one number.
MeasurementThe gates release · mandated compliance and discretionary precision/recall scored from the trace, never averaged · three reference suites · four adapters (anthropic, openai, http, mock) · file-based, single-user, not a hosted service.
tag-kit
Structured tagging primitives for human-in-the-loop annotation: scoped tags you can aggregate and score, not free-text labels.
Measurement@tag-kit/core (zero runtime deps) · @tag-kit/ui (headless React) · extracted from a real moderation app (inertial).
Collapse
A Claude Code skill-building framework: lessons, notebooks and docs compile into skills and MCP tools, linted and written atomically.
ToolingActive development. Shipped: MDX ingestor with 21 cross-stack reference lessons, notebook ingestor with admonition auto-prefill, template engine with cross-language equivalents + trigger-phrase derivation, atomic persistence, three-tier skill quality linter, and MCP server scaffold generation. One annotation can now become a skill or a working MCP tool.
Hologram
Live observability for Blender → glTF pipelines: an append-only log, a dashboard beside the assets, and an MCP surface for the agent.
ToolingOn PyPI as hologram-gltf · Python 3.10+ · MIT.
Components
all 20 primitives- Collaborative3 of 5
Written together, turn by turn.
AI Generation Scale
How much of this did a person do
- Running
Research agent
Climate policy
CompletedWriting agent
Section 2
Subagent Status
What the agent is doing right now
- HumanMostly humanCollaborativeMostly AIAI
Provenance badges
The same scale, dense enough for a table cell
Thesis
read the paperEvaluate AI on whether it assists humans without displacing them, not on whether it can finish the task alone.
Most AI systems are evaluated on whether they can complete tasks autonomously. But in deployment, they need to assist humans, not replace them. That mismatch is why 95% of enterprise AI pilots fail.
Research
all findings- A rubber stamp is also an approval: recording whether oversight was realEssay2026-08-31
- A reframe that breaks things: rebuilding eval-kit around authorizationEssay2026-08-11
- Absence passes: what a green check cannot tell youEssay2026-08-06
- Signals, not verdicts: the gate thesis appliedEssay2026-07-26
- The gate is the unit of measurementEssay2026-07-21