# humanish > Open-source (MIT) TypeScript CLI. Synthetic personas use your app in a real browser on a > local or hosted desktop. What they did lands in your repo: captures, actions, findings, and usage evidence. ## When to use humanish Reach for it when the question is "what happens when someone who is not us tries this?" and the answer needs evidence rather than a prediction: - a first-time visitor, a keyboard-only user, or a phone user has to get through a real flow (signup, first task, the thing the README promises) on a web app you maintain; - a maintainer wants a defect report they can act on: screenshots, the participant's own words, a per-task funnel, a fail-closed verify grade, and a public-safe issue draft; - a change under the UI (a probe, a prompt, a model) has to be checked against an instrument with known answers: `bench/` carries the planted-defect arms (58 of 60 defects reported over four runs, nothing invented in 15 clean runs, receipts with run ids). Not the tool for load testing, pixel diffing, or a decision about people that needs people; the limits page below says where simulated users diverge from real ones. How an agent calls it: `npx humanish init --yes`, then `npx humanish run try-live` (a real drawDB study, with a $2 estimated model-spend cap and separately billed desktop time). `npx humanish run first-run` is a synthetic evidence preview with no keys or spend; it does not study an app. Read results with `npx humanish runs --json` and `npx humanish review --run latest --json`; draft feedback with `npx humanish feedback issue --run latest --repo owner/repo`. `--json` is available where supported; `tui` is a human surface and refuses detected agent sessions. ## Install ``` npm i -D humanish @e2b/desktop npx humanish init --yes npx humanish keys set e2b npx humanish keys set openai npx humanish doctor npx humanish run try-live ``` Node 20 or newer. Install `humanish` and `@e2b/desktop` in the same project for live runs; a one-shot `npx humanish@latest` can miss the optional desktop peer. The keyless preview only needs `humanish`. - Quickstart: https://humanish.dev/docs - Your own app: https://humanish.dev/docs/your-app - Results and feedback: https://humanish.dev/docs/read-results - Paired Save-button study: https://humanish.dev/docs/save-button-study - TodoMVC keyboard repair: https://humanish.dev/docs/todomvc-edit-study - Budgets and privacy: https://humanish.dev/docs/budgets-and-privacy - Lab manifests: https://humanish.dev/docs/lab-manifests - Computer-use controls: https://humanish.dev/docs/computer-use - Local coding agents: https://humanish.dev/docs/local-agents - Observer, TUI, and JSON alternatives: https://humanish.dev/docs/review-surfaces - CLI reference: https://humanish.dev/docs/cli For coding agents: `npx skills add danielgwilson/humanish --skill humanish` ## Commands This command index is generated from the shipped CLI. Full arguments and options: https://humanish.dev/docs/cli - `humanish init`: Set up humanish/ source and .humanish/ runtime state. - `humanish doctor`: Explain project readiness and missing setup. - `humanish tui`: Human terminal for labs and runs; refuses detected agent sessions and non-TTY input/output. Agents: humanish lab list --json, humanish lab inspect --json, humanish runs --json. - `humanish telemetry`: Show or change anonymous usage collection. - `humanish telemetry status`: What is collected, and whether it is on. - `humanish telemetry enable`: Turn anonymous usage collection on. - `humanish telemetry disable`: Turn anonymous usage collection off. - `humanish keys`: Manage the user-level provider key store. - `humanish keys set`: Store one provider key in the user store (0600), prompted with hidden input. - `humanish keys unset`: Remove one key from the user store. - `humanish keys list`: List the NAMES stored in the user store. Values are never printed. - `humanish run`: Run a persona/scenario simulation or dry-run bundle. - `humanish verify`: Validate a run bundle and public-safety gates. - `humanish cleanup`: Write a resource cleanup inspection receipt. - `humanish review`: Build a review packet from verified run evidence. - `humanish analyze`: Generate evidence-linked study findings. - `humanish analyze list`: List immutable analysis versions, including failed attempts. - `humanish analyze show`: Read validated analysis and correction history. Defaults to the latest usable version. - `humanish analyze correct`: Append a human review note bound to one exact finding version; original claims remain intact. - `humanish runs`: List local Humanish runs and latest pointers. - `humanish stats`: Roll up cost, outcomes, and durations across runs. - `humanish export`: Export Observer HTML or a redacted bundle workspace. - `humanish comms`: Off-app comms surfaces. - `humanish comms providers`: List installed communication provider capabilities. No network requests. - `humanish comms connections`: Manage project-local non-secret connection profiles. A lab explicitly selects its receiving connection. - `humanish comms connections list`: Show saved connections and local credential status; does not authenticate with a provider. - `humanish comms connections add`: Save an AgentMail connection profile. Does not write a key, alter a lab or contact the provider. - `humanish comms check`: Check connection and credential presence; --online authenticates without creating inboxes. - `humanish comms configure`: Preview or save a local receiving-enabled copy of a supported lab. No provider requests. - `humanish comms recover`: Inspect interrupted email leases; --apply deletes only privately recorded resources owned by this project and account. - `humanish comms catch`: Run the adopter-hosted email catch. - `humanish runtime`: Prepare or inspect the local browser runtime. - `humanish runtime status`: Check local Docker, virtualization and the cached browser image without downloads. - `humanish runtime setup`: Download and install the local browser image. Does not start a study or use model quota. - `humanish reclaim`: Reclaim an interrupted run's sandboxes by recorded id. - `humanish watch`: Run sims, open the observer, keep the shell attached. - `humanish observe`: Follow a run's saved evidence over loopback http. - `humanish serve`: Serve the run library; optional tunnel-edge exposure. - `humanish codex`: Run Codex-native Humanish integration surfaces. - `humanish codex app-server`: Run a browser-visible Codex app-server actor surface and write redacted protocol artifacts. - `humanish lab`: List, inspect, and run Humanish lab manifests. - `humanish lab list`: List committed and ignored Humanish lab manifests. - `humanish lab inspect`: Inspect a Humanish lab manifest without running it. - `humanish lab preflight`: Check lab metadata or explicitly probe reachability. Metadata mode does not verify setup; use doctor --lab first. - `humanish lab cleanup`: Sweep stale provider resources from a crashed prior process, by provider metadata, without printing provider ids. humanish never enumerates an account by default: set HUMANISH_OSS_META_ALLOW_PROVIDER_LIST=1 to opt in for this maintainer-only sweep. - `humanish lab run`: Run a Humanish lab manifest. Same as `humanish run `, grouped under `lab`. - `humanish lab oss`: Alias: run the bundled OSS meta-lab dry-run contract. - `humanish lab oss-smoke`: Clone lightweight public OSS repos, try Humanish setup/proof, then discard clones. - `humanish feedback`: Create public-safe feedback drafts, no GitHub API. - `humanish feedback list`: List recorded feedback candidates and any saved draft. - `humanish feedback draft`: Generate a public-safe feedback draft from verified evidence. - `humanish feedback verify`: Verify the feedback draft for public issue eligibility. - `humanish feedback issue`: Print Markdown for a public GitHub issue. Does not mutate GitHub. - `humanish feedback issue-url`: Print a prefilled public issue URL. Does not mutate GitHub. `humanish tui` is for a human: Node 22+, interactive stdin and stdout. It refuses detected coding-agent sessions even with a TTY (`HUMANISH_TUI_AGENT_SESSION`), or non-TTY input/output (`HUMANISH_TUI_REQUIRES_TTY`); both return exit code 2. `--json` reports that refusal; it does not make the TUI usable by an agent. Use `humanish lab list --json`, `humanish lab inspect --json`, and `humanish runs --json` for read-only browsing. `humanish review --run latest --json` reads an existing result. `humanish lab run --json` starts a run and may incur provider spend; it is not a browsing command. `--force` is only for a person actually at the keyboard; it does not bypass TTY or Node requirements. Details: https://humanish.dev/docs/review-surfaces#for-coding-agents-and-scripts ## Credentials Choose the execution route before requesting credentials. Local browser studies use a supported Codex CLI with file-backed ChatGPT login on Linux x64 with Docker/KVM/TUN or M3+ Mac with Lima. They need neither E2B nor an OpenAI API key; participants and the separate default analyst use account quota and remote inference. Setup: https://humanish.dev/docs/your-app#study-your-running-local-app The hosted API route needs: - `E2B_API_KEY`: the hosted sandbox the persona works in. Set it with `npx humanish keys set e2b`, or `e2b auth login`. - `OPENAI_API_KEY`: the model driving the persona. Set it with `npx humanish keys set openai`. Alternatively a signed-in local agent (Codex, Claude Code) can supply the participant, but `E2B_API_KEY` is still required for that hosted desktop. Analysis defaults to the OpenAI API separately; an account participant does not change that default. `humanish doctor --lab --json` checks the selected route's setup and names the command to fix missing prerequisites without printing key values. Persisted text is scrubbed for known secret values and patterns; this is not a guarantee that all sensitive content is detected. Node 20 or newer for the CLI; the human TUI requires Node 22 or newer. ## Worked example: the Excalidraw study Run `cua-2026-08-07T17-44-48-760Z-87389419` (2026-08-07): four computer-use lanes on hosted 1920×1080 desktops drove a commit-pinned clone of Excalidraw (excalidraw/excalidraw@4872083c), an open-source virtual whiteboard. 3/4 lanes passed; one gave up; `humanish verify` ran 16/16 checks. Wall-clock 6m 06s; estimated cost ~$1.54 (rates as of 2026-08-05). Lanes and verbatim outcomes: - diagram-login-flow — persona final report: "Done" - sticky-notes — persona final report: "Done." - sketch-shapes — recorded reason (harness; the persona emitted no final message): "gave up: 8 consecutive turns with no change to the UI state" - export-drawing — persona final report: "Done" Excalidraw is the application studied; it is not a Humanish adopter or endorser. ## Credentials and sharing - Computer-use model keys stay on the host. Terminal actors receive a command-scoped runtime key by default; opt-in `execution.runtimeAuth: openai-egress` keeps the raw key outside but still permits sandbox processes to spend through the proxy. See https://humanish.dev/docs/budgets-and-privacy#store-credentials. - Evidence lands in gitignored `.humanish/` by default. Exposing Observer or sharing an export makes it accessible to others. - `humanish feedback issue` renders a draft; no GitHub API call exists on that path. - Verify fails closed: a bundle that cannot pass every gate grades `local_only` or `blocked`, never `share_ready`. ## Known failure modes The limits, cited: what the computer-use substrate cannot do yet, where simulated users diverge from real ones, what a passing `humanish verify` grade does not prove, and when to recruit real people instead. Includes first-party run ids for our own failed lanes, and the published numbers against simulated users. - https://humanish.dev/failure-modes ## Component registry humanish.dev serves a shadcn-compatible component registry at `https://humanish.dev/r/.json`. Items: `humanish-tokens` (the site's token system), `terminal-cast` (static terminal transcript), `persona-lane` (run lane with screenshot, verbatim report, honest status chip), `pinned-replay` (the scroll-pinned run walkthrough, with the Excalidraw study's metadata as demo data — screenshots not included). Install directly: ``` npx shadcn@latest add https://humanish.dev/r/pinned-replay.json ``` Or configure the namespace in `components.json` and add by name: ```json { "registries": { "@humanish": "https://humanish.dev/r/{name}.json" } } ``` ``` npx shadcn@latest add @humanish/terminal-cast ``` ## Links - GitHub: https://github.com/danielgwilson/humanish - npm: https://www.npmjs.com/package/humanish