Getting Started
This page takes you from nothing installed to a merged pull request. It covers the four ways to run harness-kit: plain, with the Jev judge, with graph engineering, and with both. Every setup uses the same commands; the options only change what happens inside some steps.
Why it is worth it
An AI agent writes code fast, and fast is where the expensive mistakes hide: the problem was never written down, "done" was never defined, the code ignores your conventions, and the review was a glance. harness-kit makes each of those a step that has to pass before the next one starts, and it writes them for you.
| Without harness-kit | With harness-kit |
|---|---|
| The ticket lives in your head or a chat thread | A PRD with the problem, the customers, and a numeric success metric |
| "Done" is whatever the agent decided | A PRP with acceptance criteria you can check one by one |
| The agent edits whatever it finds | A plan that names the files, and, with graph engineering, a gate that fails on any other file |
| Review is reading a diff and hoping | Every document scored against a rubric, retried until it passes 8.0 |
| Nobody remembers why the code is like this | Every requirement linked to the code that implements it and the test that proves it |
| Quality depends on the day | The same six steps every time, with a score history you can watch |
You spend your attention on two decisions, whether the direction is right and whether the PR is ready, instead of on typing specs and chasing conventions.
Pick a setup
You can change your mind at any time with one command, so start with the default.
| Setup | Best for | What it adds | Needs | Cost |
|---|---|---|---|---|
| Plain (default) | trying it out, small teams, most features | the full gated pipeline | Claude Code, python3, git, gh |
nothing extra |
| Jev judge | when you want a judge that is not Claude grading Claude | quality scored by Jev, a different model, in under a second per document; Claude takes over when Jev is unsure | a TypeSafe API key | about $0.04 per million tokens |
| Graph engineering | shared codebases, regulated work, anything you must audit later | requirement ids, a trace file in every PR, a scope gate, links that turn VALIDATED or STALE | nothing extra (manifest); Python 3.12, a free NVIDIA key and joern for the graph database (full) |
free |
| Graph and Jev | teams that want both a neutral judge and full traceability | both of the above | both of the above | about $0.04 per million tokens |
Install
Inside Claude Code:
/plugin marketplace add Pierry/harness-kit
/plugin install harness-kit@harness-kit
Restart Claude Code so the plugin loads. Open the repository you want to work on and run:
/harness-kit:install
It copies the agents, commands, and hooks into .claude/, adds AGENTS.md and CLAUDE.md, and asks two questions. Answer them for your setup.
| Question | Plain | Jev judge | Graph engineering | Graph and Jev |
|---|---|---|---|---|
| Eval judge | local |
jev |
local |
jev |
| Graph engineering | off |
off |
manifest or full |
manifest or full |
Restart Claude Code once more so the commands appear. Your own status line stays as it was.
Set up Jev (Jev setups only)
Create a key at console.typesafe.ai/keys and run /hk:eval jev. It asks whether the key is already in an environment variable or whether you want to paste it. A pasted key goes to .claude/settings.local.json, which git ignores; it never lands in a committed file. Restart Claude Code if you just added the key. Check it any time with python3 .claude/scripts/hk-config.py get eval.
Set up the graph (graph setups only)
/hk:graph manifest needs nothing else. For the graph database, run /hk:graph full and then:
python3 .claude/scripts/graph.py setup
python3 .claude/scripts/graph.py index-code
python3 .claude/scripts/graph.py status
/hk:graph full asks for a free NVIDIA build key from build.nvidia.com, which Graphiti uses to read your decision documents; without it the code and trace layers still work. setup creates a small virtual environment with the embedded FalkorDB and Graphiti. index-code builds the call graph of your code with Joern, once; on a large repo it takes minutes. status shows what is ready and warns if NVIDIA retired a configured model. To load decision documents, run python3 .claude/scripts/graph.py ingest docs/decisions/*.md.
Step 1: write the brief
A brief is four lines. Type it after /golden-path, or fill the brief builder, which checks each field and gives you the prompt to paste.
/golden-path
Squad: checkout
Problem: Returning guests abandon checkout when a card is declined once.
Hypothesis: If we add one-tap retry, completion rises 5 points.
Success metric: checkout completion, from 71% to 76% within 30 days
/golden-path stops for your approval after every step. If you want it to stop only twice, use /pipeline:run "<idea>" instead; it gathers context from the repo first and pauses at the same two decisions described in steps 2 and 6.
Step 2: the PRD, and your first decision
The product manager agent writes .claude/runtime/outputs/pm/prd/{feature_id}.md: problem, customers, scope, success metrics with baselines, rollout, and risks. A script checks that every section is there. Then the eval scores it on eight dimensions, each split into small yes/no checks, and below 8.0 the agent rewrites only the checks that failed.
In the plain and graph setups a fresh Claude subagent scores it. In the Jev setups Jev answers each check in one call, and if its unsure answers could flip the result, a Claude subagent decides instead. You read the PRD and approve the direction. This is the decision that matters most: a wrong problem caught here costs one rewrite, not a feature.
Step 3: the PRP
The agent turns the PRD into an engineering spec in .claude/runtime/outputs/pm/prp/{feature_id}.md, with the files to change, the patterns to follow, links to library docs, validation commands, and acceptance criteria. It searches your code with semble, repowise, or grep, whichever you have.
With graph engineering each acceptance criterion gets a stable id such as REQ-001, and trace/{feature_id}.yml is created with those requirements. From here on everything keys off the id, not the text.
Step 4: the plan
The staff engineer agent writes the plan: what changes, in which files, in what order, with risks and test cases.
With graph engineering the plan first settles older trace files, turning links from merged features into VALIDATED or STALE. Then, for each requirement, it asks for related knowledge and the code symbols most likely affected, and records each pick as a PROPOSED link with its evidence and confidence. The list of those files becomes the scope: the only files the next step may change. With full, it also sees who calls each symbol, so the blast radius goes into the risks.
Step 5: dev
The agent implements the plan in small commits and runs your linters and type checkers through the harness sensors. It follows your conventions from .claude/conventions/ when you have them.
With graph engineering each commit is recorded as an IMPLEMENTS link, and the commit hash is checked against git before it is written. Before the step ends, the scope gate compares the diff with the plan. One file outside scope fails the step. To change a file that was not planned, the agent has to add it with a written reason, which then appears in the dev summary for you to see.
Step 6: test, the PR, and your second decision
The agent runs your test suite and reports pass or fail with failures by name. With graph engineering each test that proves a requirement is recorded as VERIFIED_BY, and requirements with no test are listed as gaps.
Then it prepares the pull request: title, summary, test plan, and links. With graph engineering it validates the trace file, commits it, and adds a traceability table to the PR description: each requirement, the code it affects, its status, and the test that proves it. You approve, and the PR opens as a draft.
Step 7: merge
A monitor watches the PR and clears the pipeline when it merges. Start the next feature with a new brief. With graph engineering the next plan settles this feature's links, so what this feature proved, and what later changes broke, is visible to the next one.
What changed between the setups
| Step | Plain | Jev judge | Graph engineering | Graph and Jev |
|---|---|---|---|---|
| Scoring | Claude subagent | Jev, Claude when unsure | Claude subagent | Jev, Claude when unsure |
| PRP | criteria | criteria | criteria with REQ ids, trace file |
criteria with REQ ids, trace file |
| Plan | files to change | files to change | links with evidence, files become the scope | links with evidence, files become the scope |
| Dev | conventions and linters | conventions and linters | plus scope gate, IMPLEMENTS links | plus scope gate, IMPLEMENTS links |
| Test | report | report | plus VERIFIED_BY links | plus VERIFIED_BY links |
| PR | summary and test plan | summary and test plan | plus traceability table | plus traceability table |
Changing your mind
/hk:eval local or /hk:eval jev switches the judge. /hk:graph off, manifest, or full switches graph engineering; turning it off leaves existing trace files in git untouched. /pipeline:continue resumes a feature from where it stopped, and hk status shows where that is.
Where to go next
Golden Path for every detour, Evals and Jev and System One for how scoring works, Graph Engineering and Graph Theory for traceability and the graph, and Pipeline and Stages for what each step writes.
This page mirrors Getting Started in the wiki. Edit it there.