human(ish)

Use Humanish as a library

Run a study, read a run, bring your own participant, or score a run from TypeScript or JavaScript.

import ... from "humanish" exposes what an adopter needs for four things: run a study, read a run, bring a participant, and score a run. Everything else is reached through the humanish command.

Run a study

import { STUDY_SCHEMA, parseStudy, runStudy, type StudyResult } from "humanish";

const parsed = parseStudy({
  schema: STUDY_SCHEMA,
  id: "library-example",
  title: "Library example",
  route: "computer-use",
  subject: { source: "app-url", appUrl: "http://127.0.0.1:3000/" },
  actor: { type: "openai-computer-use", mission: "Explore the app." },
  execution: { target: "e2b-desktop", timeoutMs: 15_000 },
});
if (!parsed.ok) throw new Error(parsed.error.message);

const outcome = await runStudy(parsed.config, { cwd: process.cwd(), dryRun: true });
if (outcome.route === "computer-use") {
  const result: StudyResult<"computer-use"> = outcome.result;
  console.log(result.runId, result.ok);
}

parseStudy takes a decoded object, not YAML text: a humanish.study.v3 study. It refuses a 0.107 humanish.lab.v2 one with HUMANISH_STUDY_V2_UNSUPPORTED; humanish migrate converts the file. The StudyConfig it returns has the file's keys (route, actor, participants, caps, mode) with defaults filled, so parsed.config.actor.type is the actor and parsed.config.caps?.maxUsd the budget. A config you build yourself has the same keys.

runStudy refuses a config that still sets a StudyConfig field of humanish 0.110 (actors, a scenario object, execution.caps, subject.topology, personas, or count, lanes, roster or laneFocus on actor) with HUMANISH_STUDY_V2_UNSUPPORTED, naming each field's v3 key. It refuses a config with no actor, or a route its subject and actor do not take, with HUMANISH_STUDY_INVALID. Each refusal comes in the route's result before anything runs. runStudy's options are RunStudyOptions: env, scorer, createProvider, inProcess, prepareDesktop, onEvent, onStream, analysisSignal and onObserverReady, beside dryRun, runId, count and rerun. The library options contract lists which routes accept each one and the refusals for the rest. routeOf(config) returns the route a config takes, and StudyResult<route> is that route's result type.

onStream receives each participant's participantId, recordId (the id of its entry in run.json simulations[], such as sim-001) and streamId, and on ready its sandboxId and url.

Read a run

import { readFile } from "node:fs/promises";
import { renderObserver, verifyRun, type RunBundle } from "humanish";

const checked = await verifyRun(process.cwd(), "latest");
if (!checked.ok || !checked.bundlePath) throw new Error("the run did not verify");
const bundle = JSON.parse(await readFile(checked.bundlePath, "utf8")) as RunBundle;
console.log(bundle.review.verdict, checked.shareSafety.status);
await renderObserver(process.cwd(), checked.run, { open: false });

bundle.review.verdict is what the participants experienced. A route result's ok, also written to the run's status.json as outcome.ok, says whether the run worked as an execution; outcome.execution.failures names each failure, such as a model provider's cleanup that could not be confirmed. A run can read pass with ok: false.

This reads trusted local output. RunBundle["streams"][number] and RunBundle["simulations"][number] name a bundle's members.

Bring a participant

A participant is a ComputerUseProvider (the brain) driving a ComputerUseExecutor (the desktop or app). Pass createProvider to runStudy to replace the default brain on a computer-use study, and inProcess: { executor } with it to drive an app in this process with no desktop. An in-process run uses the same participant runner as hosted participants, so a declared caps.maxUsd or caps.maxTotalUsd applies: report usage from your provider, or the run stops at its first request with stopCause: "usage_unreported". The participant example runs one end to end:

npm install humanish
node node_modules/humanish/examples/participant/run.mjs

runComputerUseLoop is the loop underneath, for a caller that owns its own evidence; defaultRedactionHooks is the redaction it requires. createOpenAiResponsesProvider builds the default brain, and ComputerUseAdmissionLimitError is how a provider refuses a request before sending it.

Score a run

An AdapterScorerModule exports score, deriveFeedback and, on browser routes, deriveArtifacts. Pass it as runStudy's scorer, or load the same file with humanish run <study> --scorer <file>. The scorer example does both. Every scoring context carries ctx.bundle and ctx.runId.

A scorer written for one context passes through a helper that names it: scorer: browserScorer({ score: (ctx: BrowserScoringContext) => … }) for computer use and shared world, and terminalScorer(…) with TerminalProductScoringContext for terminal runs. An inline scorer, or one that narrows ctx at runtime, passes as it is.

Both contexts name the study as ctx.studyId and the run as ctx.runId. A browser scorer's context also carries ctx.route (computer-use or shared-world) and ctx.participantCount.

Removed exports

0.109.0 removed the 0.107 names that 0.108.0 kept as deprecated aliases: runLab, parseLabConfig, LAB_CONFIG_SCHEMA, LabConfig, LabEvent, LabOutcome, LabResult, LabRoute, RunLabOptions, BrowserLabScoringContext, the nine Cua* loop types and CuaAdmissionLimitError. Import runStudy, parseStudy, StudyConfig, StudyEvent, StudyOutcome, StudyResult, StudyRoute, RunStudyOptions, BrowserScoringContext, the ComputerUse* types and ComputerUseAdmissionLimitError. It also removed labId on a study's result (read studyId), simId on onStream events (read recordId), and BrowserScoringContext.backend and laneCount (read route and participantCount). A scorer reads ctx.studyId where it read ctx.labId, and a computer-use error's error.name is its ComputerUse… class name.

The route runners runCuaActorLab, runScriptedBrowserLab, runTerminalProductLab and runConcurrentSharedWorld, their option types, and the hook bag types are removed, along with the cuaHooks, scriptedHooks, terminalHooks, sharedWorldHooks and automaticAnalysis options. Use runStudy with the options above. A JavaScript caller that still passes a bag is refused with HUMANISH_STUDY_OPTION_UNSUPPORTED, naming the option to use.

0.107.0 removed runDryRun, runCuaActorSession, the route result types, LabOutcome.backend, LabBackend, selectLabBackend, the routesTo* predicates and the lab helpers that parseStudy and routeOf cover. Use runStudy and its StudyResult<route>, runComputerUseLoop, outcome.route and routeOf(config). Library options names the replacement for each.

Edit this page on GitHub