The primitive library.

Every primitive is the physical embodiment of a claim from the paper, and every specimen imports the shipped component, so the catalogue cannot drift from what the registry ships. Interactive, shadcn-compatible, copy-paste ready.

20 primitives · 17 specimens · MIT

Decision

5 specimens

The moments where the human answers. Interrupt boundaries, binary approvals, question sets, queues, and the plan the agent has not run yet.

Interrupt Cards

hitl-card

Human-in-the-loop interrupt cards rendered inline in a chat thread. Three semantic variants, each with idle, expanded, confirmed, and dismissed states, and every resolution can be undone. Click any card to expand it.

searchvariant="search"
reviewvariant="review"
writevariant="write"

Approve / Reject

approve-reject-row

The core decision row used across review, download, and notes panels. Three answers, not two: approve, reject, and can't tell, because an unresolved question is not a no. Undo returns to pending.

Verify citation accuracyIPCC 2023 · p. 12
Confirm highlighted quotePolicy Brief §3.1
Approve section for exportWriting · Section 2
Download: Carbon Pricing paperNature Climate, 2023

QA Flow

qa-flow

Multi-question approval card: single choice, multi-select, and a freeform text field. Submits to a resolved line that keeps the answers visible and can be reopened.

QA formfill out, then Continue
Preferred mechanism?
Implementation challenges?
Other notes

Batch Approval Queue

batch-queue

Sequential approve-reject flow across mixed agent items. Auto-advances to the next item, can step back one, and resolves to a summary state.

Kitchen sink batch5 items
  1. Search: carbon pricing 2024
  2. Write: Section 2 introduction
  3. Research: IPCC AR6 findings
  4. QA: Verify citation accuracy
  5. Read: eu-ets.europa.eu

Editable Plan

editable-plan

The multi-step plan the human can rename, reorder, add to, or delete from before the agent executes it. Steps marked locked cannot be removed.

Research planedit any unlocked step
Draft research summaryReorder, edit, or remove steps before the agent runs

Agent state

4 specimens

What the agent is doing and what it is about to do. Execution states, the reasoning trace, the pending tool call, and the context it was handed.

Subagent Status

subagent-status-card

Six discrete agent execution states. The running state animates. Use it in any card that wraps an in-progress agentic task.

idlestatus="idle"

Research Agent

Climate Policy workspace

Idle
runningstatus="running"

Research Agent

Climate Policy workspace

Running
completedstatus="completed"

Research Agent

Climate Policy workspace

Completed
errorstatus="error"

Research Agent

Climate Policy workspace

Error
skippedstatus="skipped"

Research Agent

Climate Policy workspace

Skipped
cancelledstatus="cancelled"

Research Agent

Climate Policy workspace

Cancelled

MiniTrace

mini-trace

Step-by-step thought, action, result renderer; each step collapses to reveal its detail. A visible implementation of the supporting-facts requirement from §3.3 of the paper.

Search trace3 steps · click a row

Tool Call Preview

tool-call-preview

The tool call the agent wants to make: name, arguments, optional rationale and signals, shown before it runs so the human can approve or reject. Pairs with the gates layer for confidence, cost, and scope checks.

Outbound emailexpand Arguments
send_email()tool call

Drafted reply to the client thread; confirming high-stakes outbound before sending.

86% confidence$0.0012write:emailsend:external

Context Chips

context-chips

Pill chips for the context attached to an agent run: notes, files, URLs. Removable, with overflow truncation built in.

Context stripclick × to remove
  • AR6 temperature finding
  • IPCC AR6 Synthesis.pdf
  • eu-ets.europa.eu
  • Price corridor note
  • Carbon Markets 2024.pdf

Evidence

4 specimens

What the agent found, and where it came from. Ranked results, source-backed claims, and the exact text a proposed edit would change.

Search Result Cards

search-result-card

Ranked result cards with metadata, snippet, and a relevance bar. The relevance figure is a signal for the human to weigh, not a verdict the agent has already acted on.

Result #1Nature Climate Change, 2023 · 97%

Carbon Pricing Mechanisms and Emissions Outcomes

Stavins, R., Stowe, R., Comstock, M., Nature Climate Change, 2023412 citations

Empirical analysis of 40 carbon pricing schemes across 30 jurisdictions reveals that price levels above $50/tCO₂ are associated with meaningful emissions reductions in the power sector...

97%
Result #2arXiv, 2024 · 91%

Just Transition Frameworks in Coal-Dependent Regions

Newell, P., Mulvaney, D., arXiv, 202487 citations

A systematic review of 28 coal phase-out programmes finds that regions with co-designed transition plans achieved 34% higher re-employment rates within 24 months...

91%
Result #3Science, 2023 · 88%

Net Zero Pledges: Verification and Accountability

Höhne, N., Gidden, M., den Elzen, M., Science, 2023301 citations

Of 196 national net zero pledges examined, only 31% include robust interim milestones and independent verification mechanisms aligned with 1.5°C pathways...

88%
Result #4Energy Economics, 2024 · 84%

EU Emissions Trading System Post-Reform Price Dynamics

Borghesi, S., Flori, A., Energy Economics, 2024156 citations

The Market Stability Reserve mechanism introduced in 2019 has demonstrably reduced permit surplus, contributing to a tripling of EUA prices from 2018 to 2023...

84%

Citation Result

citation-result

A single source-backed citation: the claim on top, the source attribution below, an expandable supporting quote, and an optional confidence badge. Verify, reject, or can't tell.

Cited claimexpand the supporting quote

Roughly 95% of enterprise generative-AI pilots fail to reach production deployment.

Challapally et al.2025MIT NANDA78% confidence
The GenAI Divide: State of AI in Business 2025

Evidence Pointer

evidence-pointer

Where a claim is grounded, not merely that it is. One row per pointer with the source, the locator in human units, and the excerpt. Sources the agent consulted and drew nothing from are listed too, so silence never reads as safety.

Located claimtwo pointers, two sources not assessed
Evidence2 located · 2 not assessed

The comment addresses the recipient by name and repeats a threat made in an earlier post.

  • Comment #8813characters 42–97
    …you know exactly where I'll be waiting for you, Dana.
  • Earlier post #8801characters 0–61open
    I'll be waiting outside after the meeting.

Not assessed: Direct messages between the accounts, Reports from other users

Diff Result

diff-result

Before and after for a proposed text or code edit, with per-hunk strips. Drop it into any agent loop where the human should see exactly what will change before it lands.

Markdown rewriteApply edit to confirm
Tighten introduction paragraphRemoves hedging, names the thesis directly
markdown
@ line 1
It seems like the central question of this paper might be how AI systems should be evaluated when they are deployed in collaborative settings.
This paper argues that AI systems must be evaluated against collaborative performance, not autonomous task completion.

Composed

2 specimens

Whole task surfaces built from the primitives above. The shape a real agent panel takes once the parts are assembled.

Writing Agent

writing-agent

A compound widget for a draft in progress: title, target section, word range, evidence notes, and the same six status states the subagent card uses.

Write doc agentclick a status chip to cycle
Write Doc Agent
Idle
Title
Climate Policy Analysis
Target
Section 2
Word range
400–600

Evidence notes

  • AR6 temperature overshoot
  • Price corridor $50/tCO₂
  • EU ETS reform outcomes

Research Agent

research-agent

Three operating modes for a long-running research task: create a new session, follow up on an existing one, or read a single URL.

Research agentswitch modes, top right
Research Agent
Profile
Academic, Climate Policy
Engine
Semantic Scholar + Web
Depth
Deep (5 hops)

Scales & palette

2 specimens

How much of this did a person do. One five-point scale from Human to AI in four densities, slider to inline pill, plus the five accents and four approval badges the kit draws from.

AI Generation Scale

four densities

One question, how much of this did a person do, answered on a five-point scale from Human to AI. Every density shows the same thing: a track filled to the current level, revealing the spectrum from emerald to rose as it goes, and the level in plain words. Colour is the cue; the words carry the meaning. Labels sit on foreground and muted-foreground, never on colour.

Sliderai-generation-slider
Collaborative3 of 5
section 2 draft

Written together, turn by turn.

Reach for this when the person sets the value. The readout says the level and what it means, and the track fills as far as the thumb, revealing the spectrum. Drag it, tap either end, or focus it and use the arrow keys, Home and End.

Meterai-generation-meter
HumanMostly humanCollaborativeMostly AIAI
compact

Reach for this to show provenance in a list row or a header without inviting interaction. Read-only by design: one image element with no focusable children, so fifty rows do not become fifty tab stops.

Badgeai-generation-badge
Mostly AI, 4 of 5interactive
HumanMostly humanCollaborativeMostly AIAI

Reach for this in a table cell or a queue row where even the meter is too much furniture. Static by default; given an action handler it grows ‹ › steppers with real 24px targets that go inert at the ends of the scale without dropping keyboard focus.

Segmentedai-generation-scale

CollaborativeWritten together, turn by turn.

The explicit form, for a settings panel or a form where every option should be visible and directly tappable. Five segments in one pill, the chosen one raised, each carrying its colour so the order reads before the words do. It needs about 400px; below that, use the slider.

Shared Primitives

shared-primitives

The palette the rest of the kit draws from. Five accents, each with one job; the four approval badges, because can't tell is not no; and the decision row that sets them.

Shared paletteinteractive

Accents

colour is punctuation, never a background for text

  • VioletSearch and evidence
  • AmberReview, hold, can't tell
  • BlueWriting and running
  • EmeraldApproved
  • RoseRejected

Approval badges

four states, because can't tell is not no

PendingApprovedRejectedCouldn't tell

Decision rows

undo returns to pending

Item A
Item B
Item C

Each primitive installs on its own through the shadcn CLI. No fork, no wrapper SDK. The registry page carries the exact command and dependency chain for every item.