human(ish)

Study file reference

Every field a humanish.study.v3 file can set, with its type, default and meaning, plus persona files and scripted scenarios.

A study file is YAML with schema: humanish.study.v3. This page lists every field the parser accepts. How a study works explains the terms on one annotated file, and the own-app guide builds a complete study for your app.

Keep study files in one of three folders

humanish/studies/*.yaml          committed studies anyone who clones the project can run
.humanish/studies/*.yaml         ignored local studies
.humanish/local/studies/*.yaml   ignored private or machine-specific studies

humanish run <id> finds a study by its id in these folders. Private repository targets, preview URLs and local variants belong in the ignored folders, and you can run any study file by path:

npx humanish watch .humanish/studies/local-dogfood.yaml --dotenv .humanish/local/provider.env
npx humanish run .humanish/studies/local-dogfood.yaml --json --no-open

--dotenv loads values into the humanish process, and the names a study lists in subject.env are then passed to its sandbox. The run records those names and never their values. Persisted text is scrubbed for known values and secret patterns, with the coverage limits described under verify before sharing.

Check a study before you run it

npx humanish study list --json
npx humanish study show first-run --json
npx humanish study check first-run --json

None of the three creates a desktop or calls a model. study show prints the resolved study, including the persona text each participant receives. study check parses the file and reports problems with code HUMANISH_STUDY_INVALID:

  • A field the parser does not know fails, and the message lists the known fields at that level and suggests the closest one. A misspelled runtme in a participant entry fails with that entry's index.
  • A field the study's route does not read fails, and the message names the route, such as "route: preview reads no caps; remove the block."
  • The route must match the subject, execution.target and the actor. When they disagree, the message names the route they take.

The tables below give each field's type and its default when you leave it out. "required" means the study fails without it.

Look up a top-level field

FieldTypeDefaultMeaning
schemastringrequiredhumanish.study.v3. humanish refuses a humanish.lab.v2 file, and humanish migrate converts it.
idstringrequiredThe study's name on the command line. It starts with a letter or digit, followed by letters, digits, _, . or -.
titlestringnoneOne line that study list, study show and the TUI print.
descriptionstringnoneLonger text that study show and the TUI print.
routepreview, computer-use, shared-world, scripted or terminalrequiredHow the study runs. How a study works describes each one.
modedry-run or livedry-runlive runs participants. A dry run writes a synthetic run with no browser, model, key or spend. --dry-run on the command line forces a dry run.
subjectmappingrequiredThe product under study. Subject fields.
actormappingrequiredWhat drives every participant, and the task. Actor fields.
participantsnumber, mapping or list4 on preview, 1 on computer-use; required on shared-worldWho takes part. Participant fields.
surfaces[desktop] or [desktop, mobile][desktop]scripted only: the viewports the scenario replays on.
capsmappingnoneSpend and time limits. Each route reads its own keys. Caps.
executionmappingnoneWhere participants run and for how long. Execution fields.
scenariostringrequired on scriptedscripted only: a scenario id from humanish/scenarios/, or a path to a scenario file. Scripted scenarios.
policiesmappingevery policy falsePermissions the study grants. Policies.
reviewmappinganalysis onWhat runs after the participants finish. Review fields.
defaultsmappingnonedefaults.open (boolean): whether run and watch open the Observer when the run finishes. --open and --no-open override it. Without it, only watch in a terminal opens.
commsmappingnoneAn inbox each participant can read, for email-gated flows. Comms fields.

Declare the product in subject

subject.source says what kind of product the study runs on, and each source takes its own fields. A field set on a source that cannot use it fails at parse.

subject.sourceThe productRoutes
app-urlAn app that is already running at subject.appUrl.computer-use, shared-world, scripted
cloneRepositories humanish clones, builds and serves inside each hosted desktop.computer-use, shared-world, scripted
local-treeYour working tree, packed and served inside each hosted desktop. Gitignored files, .env* files and other secret-shaped files are left out.computer-use, shared-world
desktop-cliA CLI the participant uses in a terminal window on a hosted desktop.computer-use
terminal-productA CLI an agent studies in a terminal sandbox, starting from its public pages.terminal
local-appA dev server on your machine that your own code drives through the library. The CLI refuses it.computer-use
this-repoNothing runs: the preview writes a sample run.preview
FieldTypeDefaultMeaning
subject.sourceone of the seven aboverequiredWhat kind of product this is.
subject.appUrlhttp(s) URLrequired on app-url and local-appThe address participants open. A URL other than 127.0.0.1 or localhost needs policies.allowPublicTargets: true. On local-app it must be loopback. A user name, a password or a credential parameter such as token= in it is refused.
subject.reposlist of owner/reporequired on cloneThe repositories to clone.
subject.clonemappingnoneclone only. subject.clone.depth: git clone depth, default 1. subject.clone.keep: keep a failed participant's sandbox for debugging, default false. subject.clone.fanout is read by no route and refused.
subject.localTreemappingnonelocal-tree only. exclude: extra paths or file names to leave out of the archive. keep: keep a failed participant's sandbox, default false. maxArchiveBytes: the upload limit, default 256 MiB.
subject.servemappingrequired on clone and local-treeHow to start the app inside the sandbox (below).
subject.envlist of variable namesnoneclone and local-tree: names whose values humanish reads from your environment or --dotenv and passes to the app. The run records the names only.
subject.envValuesmapping of name to stringnoneclone and local-tree: non-secret values committed with the study, such as a public base URL or a feature flag. They are recorded in the run, so a value that looks like a secret or a local path fails at parse.
subject.statemappingnoneclone and local-tree: seed steps, external state and shared-world checkpoints (below).
subject.productmappingrequired on terminal-product and desktop-cliThe CLI under study (below).
subject.publicTarget{ owner, authorized: true }noneapp-url on shared-world only, and required there: your statement that you own or operate the public deployment. owner is a public-safe label such as owner/repo, and it is recorded in the run.
subject.exposuresyntheticnoneYour statement that the app, served at a public sandbox URL during the run, holds only synthetic data. Required on shared-world with clone or local-tree, and on scripted with clone.

Start a cloned app with subject.serve

FieldTypeDefaultMeaning
subject.serve.installshell commandnoneRuns once before the build, such as npm ci.
subject.serve.installTimeoutMsmilliseconds600000The install step's time limit. Monorepos can need more.
subject.serve.buildshell commandnoneRuns once before the app starts.
subject.serve.buildTimeoutMsmilliseconds600000The build step's time limit.
subject.serve.startshell commandrequiredThe long-running command that serves the app. A shared-world study's command must bind 0.0.0.0.
subject.serve.urlloopback URLrequiredThe address humanish waits on and the participant opens, such as http://127.0.0.1:3000/.
subject.serve.readyTimeoutMsmilliseconds180000How long the app has to answer at url after it starts.

The commands run inside the sandbox with the same trust as your own scripts. The run records a digest of each one.

Seed the app's state with subject.state

FieldTypeDefaultMeaning
subject.state.seedlist of stepsnoneCommands that prepare data, in order. Each step has name (lowercase letters, digits and -, at most 40 characters), command, when (before-build, before-start or after-ready; default before-start) and timeoutMs (default 300000).
subject.state.externallist of variable namesnoneVariables in subject.env that point at state the study does not control, such as a shared database. The run records its state as unpinned.
subject.state.checkpointlist of { name, command, redact }noneshared-world only, and required on a clone or local-tree plane: read-only commands run before and after each participant's turn. Only a digest of their output is kept; redact lists extra values to remove first.

Describe a CLI with subject.product

FieldTypeDefaultMeaning
subject.product.namestringrequiredThe product's public-safe name, recorded in the run.
subject.product.publicSurfaceslist of http(s) URLsrequiredThe pages the participant starts from, such as docs or an llms.txt.
subject.product.installshell commandnoneInstalls the product before the participant starts. Without it, installing is part of the participant's task.
subject.product.workdirpaththe home directorydesktop-cli only: the directory the participant's terminal opens in.
subject.product.uploadproject-relative pathnoneterminal-product only: a local file, such as an unpublished build, copied into the sandbox before install and exposed as $HUMANISH_PRODUCT_UPLOAD.

Choose the actor and its task in actor

A study has one actor. It drives every participant, and each participant entry can override some of its fields.

FieldTypeDefaultMeaning
actor.typeopenai-computer-use, local-agent, scripted-browser, codex-exec, or a label on previewrequiredWhat drives each participant. The actor table says what each needs.
actor.missionstringnoneThe brief every participant receives, written the way you would brief a person. Read on computer-use, shared-world and terminal.
actor.personapersona idnoneThe persona file for every participant without its own. With none, a computer-use participant gets the id cua-operator and no persona text.
actor.taskslist of { id, goal, success }nonecomputer-use only: discrete tasks added to the mission. The participant sees each goal. success is a stop rule, in the stopWhen shape, that marks the task done; the participant never sees it.
actor.modelstringgpt-5.6-sol; gpt-6-astra in a local browser; the agent's own on a hosted local-agentThe model name sent to the provider. A cap on a model humanish has no price for is refused before the run.
actor.maxOutputTokenspositive integernoneopenai-computer-use only: the most output tokens, reasoning included, in each model reply. The first request uses at most 1024. A capped study without it stops on a stalled request.
actor.localAgentcodex or claudecodexlocal-agent only: which signed-in coding agent drives the participant.
actor.reasoningEffortnone, minimal, low, medium, high, xhigh or maxthe provider's default; low in a local browserHow hard the model is asked to think each turn. A level the model does not accept fails on the first turn. The run records the level it sent.
actor.stopWhen{ any: [rules] }noneEnds a computer-use participant as soon as any rule matches what it sees (below).
actor.dwellmappingnoneA window in which humanish only watches: it takes no action and requests no model turn (below).

A stopWhen rule sets at least one of urlIncludes (text in the URL), urlPathEquals (an exact path starting with /), textIncludes (text on the page) and appStatePathEquals ({ path, equals }, a value in the app state a custom executor reports). An optional id names the rule in the run.

A dwell window has ms, its length from 1000 to 3600000 milliseconds, and is required. everyMs is how often a frame is captured, default 10000. then is continue, the default, or stop to end the session after the window. when is a stopWhen-shaped condition that opens the window; without it, the window opens after the first observation.

List the participants in participants

participants takes the forms its route accepts:

RouteFormsDefault
previewa count4
computer-usea count, { count, instruction } for identical participants, or a list of entries1
shared-worlda list of at least two entriesrequired
scriptedrefused; set surfaces instead
terminalrefused; one agent runs

A computer-use study runs at most 16 participants, each on its own desktop. A local-app subject takes no list. --count on the command line overrides a count. Each list entry accepts these fields, and an unknown one fails with the entry's index:

FieldTypeDefaultMeaning
participants[].idstring, at most 40 characterslane-01, lane-02, …; role-01, … on shared-worldNames the participant and its evidence folders. Ids are unique.
participants[].countpositive integernoneMakes the entry a group of that many participants, <id>-01 to <id>-NN, which share its other fields. A group needs an id of at most 37 characters.
participants[].personapersona idactor.personaThis participant's persona file.
participants[].instructionstringnoneText added to the mission for this participant. In the { count, instruction } form, participants.instruction goes to every participant.
participants[].devicescreen presetexecution.desktop.deviceThis participant's screen preset. It cannot be combined with execution.desktop.resolution.
participants[].reasoningEfforteffort levelactor.reasoningEffortThis participant's reasoning effort.
participants[].stopWhen{ any: [rules] }actor.stopWhenThis participant's stop rules.
participants[].dwellmappingactor.dwellThis participant's observation window.
participants[].targetabsolute http(s) URLsubject.appUrlcomputer-use with app-url only: the URL this participant opens. A non-loopback one needs policies.allowPublicTargets: true.
participants[].entrypathsubject.serve.urlshared-world only: a path on the served app where this participant starts.
participants[].hostbooleanfalseshared-world with app-url only: marks the one participant who creates the shared session, such as a game lobby, that the others join.
participants[].actorTypelabelnoneA free-form label recorded in the run for grouping participants, such as viewer. It does not choose the actor.
participants[].surfacelabelnoneA free-form label for the product area a participant starts from.
participants[].caseGrouplabelnoneA free-form label that ties participants to one shared case, account or work item.

Limit spend and time with caps

Every route reads the same top-level caps block, and each route refuses the keys it does not read. The caps act on humanish's estimates. They are not provider billing limits, and what a study costs explains what they leave out.

FieldTypeDefaultMeaning
caps.maxUsdnumber, 0 or morenonecomputer-use and shared-world: stop a participant once its estimated model spend passes this. terminal: required for a live run, and it must be 0 (below).
caps.maxTotalUsdnumber, 0 or morenonecomputer-use and shared-world: stop every participant once the study's estimated model spend passes this.
caps.maxJobsnumber, 0 or morenoneterminal: the billable product jobs the agent may trigger, checked after the session against measured job counts.
caps.maxMinutesnumber above 0required for a live terminal runterminal: the agent command's time limit. A live run refuses more than 49, or 44 with subject.product.install, because the sandbox lives at most an hour.

preview and scripted take no caps.

Cap a terminal study

humanish has no cost source for a terminal study: Codex reports no provider cost, and humanish measures no product, media or payment spend. A positive caps.maxUsd could never trip, so a live terminal run refuses it with HUMANISH_TERMINAL_UNPRICED_CAP. Set caps.maxUsd: 0 and bound the run with caps.maxMinutes. caps.maxJobs has no measured job count to check for the same reason. Adopter-supplied cost sources are planned in issue 347.

Say where participants run in execution

FieldTypeDefaultMeaning
execution.targete2b-desktop, local or e2b-terminale2b-desktop on desktop-cli, e2b-terminal on terminal-product, local on a scripted app-urlWhere participants run. e2b-desktop is a hosted E2B desktop. local is a disposable VM on your machine for a computer-use app-url study, or this machine's browser for a scripted one. clone and local-tree need e2b-desktop, and a computer-use app-url study must set local or e2b-desktop.
execution.timeoutMsmillisecondssee belowEach participant's session limit, on computer-use, shared-world and scripted. terminal refuses it: caps.maxMinutes is its limit.
execution.concurrencypositive integerevery participant at oncecomputer-use and shared-world: how many participants run at the same time. shared-world needs at least 2.
execution.desktopmappingnoneThe desktop's screen, browser, image, media and recording (below).
execution.egressAllowlist of host namesunrestrictedterminal: the only hosts the sandbox may reach. *.example.com matches subdomains at any depth, so list the apex domain separately. An empty list fails at parse. Any other route ignores the field with a warning, since its browser desktops can reach any site, and 0.112.0 refuses it there.
execution.runtimeAuthopenai-env or openai-egressopenai-egressterminal: how the agent gets its OpenAI key. openai-egress keeps it in an E2B proxy rule, and sandbox processes can still spend through that proxy. openai-env passes the key to the agent command, where the agent and its child processes can read it.
execution.runtime{ version }the latest releaseterminal: an exact Codex version to install, such as 0.153.3. Without it, humanish resolves the latest release once and records it.
execution.terminalmapping{ transport: exec-stream, stdin: disabled }terminal: transport takes only exec-stream, and stdin takes disabled or planned. stdin: sent fails at parse.
execution.completionTimeoutMsmillisecondsnoneNo route reads it, so a study that sets it fails at parse.

Without execution.timeoutMs, a session gets 30 minutes on an app-url subject. On clone and local-tree it gets what the 60-minute sandbox leaves after the clone, build and seed steps: 20 minutes for a clone with no seed steps, and never less than 5. A local browser session defaults to 20 minutes, which is also its maximum, and a scripted session to 5 minutes.

Set up the desktop with execution.desktop

FieldTypeDefaultMeaning
execution.desktop.devicescreen presetdesktopComputer-use participants: mobile (414×896), small-mobile (360×740), narrow-mobile (320×700), tablet (820×1180), desktop (1440×950) or wide (1920×1080).
execution.desktop.resolution[width, height]none; 960×720 in a local browserA raw screen size that overrides device.
execution.desktop.browserdefault, chrome, chromium or firefoxdefaultThe browser on a hosted computer-use desktop. A named browser that fails to launch fails the participant.
execution.desktop.templatestringE2B's stock desktop imageA custom E2B template name or id, for an app that needs runtimes the stock image lacks. The run records it.
execution.desktop.sandboxTimeoutMsmillisecondsderived from the sessioncomputer-use: the hosted sandbox's own lifetime. E2B allows at most 60 minutes.
execution.desktop.fidelitymappingnoneMobile emulation for hosted Chromium on a mobile preset: mobileEmulation (required), deviceScaleFactor (the preset's), touch (true) and userAgent (an iPhone Safari string).
execution.desktop.mediamappingnonecamera.source (synthetic, or a .y4m file) and microphone.source (speech, with a Codex local-agent only). Participant media has the limits.
execution.desktop.recording{ audio }noneKeep a screen video of each computer-use participant on its own desktop; audio: true adds sound. Desktop recording.
execution.desktop.codexAppServerbooleannoneNo route reads it, so a study that sets it fails at parse.

Grant permissions in policies

Each policy is false until the study sets it to an unquoted true.

FieldDefaultMeaning
policies.allowPublicTargetsfalseLets subject.appUrl or a participant's target be a non-loopback URL you own, such as a preview deployment. With more than one participant it is refused unless each participant has its own target; route: shared-world puts several on one public app.
policies.redactScreenshotsfalseBlurs and downscales screenshots when they are saved. The model still sees full frames.
policies.redactReposfalse; true when the clone authenticates with GITHUB_TOKENReplaces repository names in saved evidence.
policies.mediaPermissionpromptHow the browser's camera and microphone prompt is answered. prompt leaves the real dialog to the participant, and granted accepts it in advance.
policies.allowPrivateRepoAccess, policies.allowProviderCredentials, policies.allowPaymentCredentials, policies.allowGitHubMutationfalseterminal: recorded in the run as declared intent. No route passes these credentials to the agent.

Configure analysis in review

FieldTypeDefaultMeaning
review.analysisfalse or mappingon, with OpenAI and a $3 limitThe separate model reading of a finished live run. false turns it off. A default analysis is skipped without OPENAI_API_KEY or when its estimate is over the limit. One the study declares fails when its estimate is over its limit.
review.scorer{ ref }noneA project-relative path to a scorer module (.mjs) that exports score, deriveFeedback or deriveArtifacts. It runs as code, so review it as code. scripted refuses it.
review.scoring, review.milestones, review.vocabularystringnoneNo route reads them, so a study that sets one fails at parse.

review.analysis as a mapping takes these keys:

KeyDefaultMeaning
provideropenai; codex for a Codex local-agent in a local browseropenai uses OPENAI_API_KEY. codex uses your signed-in Codex account and takes no dollar or token limit.
maxCostUsd3openai: the analysis is refused before it starts when its estimate is over this. It is not a billing limit.
modelgpt-6-astraThe analysis model.
questionnoneA question for the analysis to answer, up to 4000 characters.
timeoutMs600000The analysis time limit, at most 600000.
maxOutputTokens16384openai: the reply's token limit, at most 32768.

Read results shows the findings, and humanish analyze runs an analysis on demand.

Give participants an inbox in comms

comms.email gives each participant an inbox for email-gated flows such as a signup link. By default humanish captures the mail your app sends inside the sandbox, which needs a clone or local-tree subject. Email capture and real email receiving have complete examples.

FieldTypeDefaultMeaning
comms.emailmappingnoneThe inbox settings below.
comms.email.kindfakefakeA captured inbox. A real inbox sets connection and leaves kind out.
comms.email.injectEnvvariable namerequired unless smtp or external is setThe variable humanish sets to its catch's URL. Your app reads it as its email API's base URL.
comms.email.portport number8025The catch's port inside the sandbox.
comms.email.smtpmappingnoneSMTP capture: hostEnv and portEnv (required, the variables your app reads), port (default 2525), and userEnv, user, passwordEnv and password for apps that require credentials.
comms.email.linkOriginhttp(s) originnoneThe origin your app writes into email links, when it differs from the served URL.
comms.email.recipientslist of { lane, address }<participant id>@example.test for each oneEach participant's address. lane holds the participant id.
comms.email.externalmappingnoneA catch you run yourself with humanish comms catch, for app-url subjects: catchBaseUrl (required), inboxBaseUrl (defaults to catchBaseUrl) and authTokenEnv.
comms.email.connectionconnection namenoneA saved real-inbox connection from the TUI's Connections screen. It cannot be combined with the capture fields above.
comms.email.allowedOriginslist of http(s) originsnoneWith connection: up to 16 more exact origins that a real inbox's links may open.

Describe a participant in a persona file

A persona file says who a participant is. humanish looks for <id>.yaml or <id>.yml in humanish/personas/, then in the ignored .humanish/local/personas/, and a committed file wins. A persona id uses letters, digits, _ and -. humanish init writes two examples:

humanish/personas/skeptical-power-user.yaml
schema: humanish.persona.v1
id: skeptical-power-user
name: Skeptical Power User
summary: A privacy-safe experienced user looking for speed, reversibility, and clear proof.
traits:
  patience: low
  technical_confidence: high
  accessibility_needs: keyboard_first
constraints:
  - Do not use real personal data.
  - Prefer synthetic fixture inputs.
  - Flag unclear recovery paths.
FieldTypeDefaultMeaning
schemastringnonehumanish.persona.v1.
idstringthe id the study namesThe persona's id, at most 120 characters.
namestringthe id in title caseWhat the participant is called, at most 120 characters.
summarystringnoneOne line about the person. Text past 280 characters is cut with a warning.
backgroundstringnoneWho the person is, in as many paragraphs as you need, up to 32 KiB of UTF-8. A larger or empty background stops the run before it starts.
traitsmappingsee belowBehavior directives, below.
traits.patiencelow, medium or highmedium without a backgroundHow much friction the participant works through before stopping.
traits.technical_confidencelow, medium or highmedium without a backgroundWhether the participant sticks to visible controls or uses shortcuts and advanced options.
traits.accessibility_needstext, at most 80 charactersnoneA need the participant works within. none, none declared and not applicable mean none.
constraintslist of stringsnoneRules the participant is told to honor. The first 8 are used, each up to 160 characters.

With a background, a trait you leave out sends no directive, and a file that sets both a background and patience or technical_confidence gets a warning to check them for conflicts. Any other field or trait warns and is not sent. An invalid level warns and falls back to medium, or to no directive when there is a background.

Each trait becomes one sentence of the participant's instructions:

TraitValueSentence the participant receives
patiencelowYou are impatient: if you hit repeated friction or stop making progress toward your goal, you are likely to stop. Describe what you actually encountered.
patiencemediumYou have moderate patience: you will work through some friction, but may stop if further effort no longer seems worthwhile.
patiencehighYou are determined: you are willing to spend effort on recovery when the goal matters to you, but may stop at an unrecoverable dead-end.
technical_confidencelowYou are not technically confident and generally rely on familiar, visible controls. Describe confusion only when you actually encounter it.
technical_confidencemediumYou have moderate technical confidence and generally use straightforward paths through the product.
technical_confidencehighYou are technically confident and comfortable with keyboard shortcuts and advanced options on the surfaces this product gives you.
accessibility_needscontains keyboard_firstYou prefer keyboard navigation, but can use a pointer when needed. Describe any difficulty you encounter switching between them.
accessibility_needscontains keyboard_onlyYou can only use the keyboard. If an essential control has no discoverable keyboard path, describe where you could not proceed.
accessibility_needsmentions terminal or outputYou rely on clear terminal output. Describe any difficulty understanding the output you actually encounter.
accessibility_needsanything elseYour accessibility requirement is: (your text). Use the available interface within this requirement and describe any difficulty you encounter.

These sentences are instructions. humanish does not stop a keyboard_only participant from using the mouse, and a persona certifies nothing about accessibility. npx humanish study show <study> --json prints the persona text each participant receives, and the run records it.

Replay scripted browser steps

A scripted study replays committed browser steps against an app you run on a loopback URL. The steps come from the scenario named by scenario and make no model requests. Browser steps are public-safe source, so use synthetic fixture values and committed relative app paths only.

humanish/scenarios/todo-onboarding.yaml
schema: humanish.scenario.v1
id: todo-onboarding
title: Todo onboarding
persona: synthetic-new-user
goal: Create the first synthetic todo and verify the list updates.
mode: browser
browser:
  startPath: /
  steps:
    - id: open-home
      label: Open the todo app
      action: goto
      path: /
      expect:
        text: Add todo
    - id: enter-todo
      label: Enter synthetic todo text
      action: fill
      selector: input[name="todo"]
      value: Synthetic onboarding task
    - id: create-todo
      label: Create the todo
      action: click
      selector: button[type="submit"]
      expect:
        text: Synthetic onboarding task
        stateChanged: true

Supported actions are goto, fill, click, press, assertText, waitForText, and waitForSelector. A press step needs key, in Playwright's key syntax (Enter, Tab, Control+A). With a selector it focuses that element and presses the key; without one it presses the key on the focused element. Supported expectations are text, selectorVisible, urlIncludes, and stateChanged.

Save the scenario as humanish/scenarios/todo-onboarding.yaml and point a study at it. humanish/studies/scripted-demo.yaml is a complete example.

humanish/studies/todo-onboarding.yaml
schema: humanish.study.v3
id: todo-onboarding
route: scripted
mode: live
subject:
  source: app-url
  appUrl: http://127.0.0.1:3000
actor:
  type: scripted-browser
surfaces: [desktop, mobile]
scenario: todo-onboarding
review:
  analysis: false

surfaces: [desktop, mobile] adds a mobile surface to the desktop one. Without mode: live the study is a dry run: it validates the steps and writes a contract bundle without opening a browser. Post-run analysis is separate from the steps: when OPENAI_API_KEY is set it runs by default, and it is refused when its estimate exceeds $3. analysis: false turns it off, so the study makes no model requests.

A live run drives a local Chrome or Chromium against your app, so start the app on http://127.0.0.1:3000 first. If humanish finds no browser, the run stops with HUMANISH_SCRIPTED_BROWSER_MISSING; set HUMANISH_BROWSER_COMMAND to the browser binary. Chrome keeps a Unix socket under TMPDIR, and a socket path holds at most 107 bytes, so when TMPDIR is longer than about 60 bytes humanish launches the browser with TMPDIR=/tmp. --dry-run checks the steps without either:

npx humanish run todo-onboarding --dry-run --json --no-open
npx humanish run todo-onboarding --json --no-open
npx humanish verify --run latest --json

Traces are stored as JSON under .humanish/runs/<run>/traces/ and summarized in the Observer.

Move files from 0.107

Before 0.108.0 a study was a humanish.lab.v2 file under a labs/ folder. humanish no longer runs those files or reads the labs/ folders. Running one by name fails with the command that fixes it:

humanish run failed: humanish/labs/old-study.yaml is a humanish.lab.v2 file in humanish/labs/, which humanish no longer reads. Run humanish migrate humanish/labs/old-study.yaml to convert it and move it to humanish/studies/.
code: HUMANISH_STUDY_V2_UNSUPPORTED

A v3 file left in a labs/ folder fails with HUMANISH_STUDY_RETIRED_DIRECTORY, and its message names the matching studies/ folder to move it to; migrate skips v3 files. humanish study list --json lists each refused file in retired, with its code and message. Convert every v2 file at once:

npx humanish migrate --dry-run
npx humanish migrate

--dry-run lists each file's destination and every key it moves or drops, with the dropped values, and writes nothing. migrate then rewrites each v2 file as v3 and moves it from labs/ to the matching studies/ folder; a v2 file outside labs/ is rewritten in place. Comments move with their keys. A key the v2 file declared but its route never read is dropped and reported, and a file that uses YAML anchors is refused with the line number.

Edit this page on GitHub