Use Humanish as a library
Run a study, read a run, bring your own participant, or score a run from TypeScript or JavaScript.
import ... from "humanish" exposes what an adopter needs for four things: run a study, read a run,
bring a participant, and score a run. Everything else is reached through the humanish command.
Run a study
import { STUDY_SCHEMA, parseStudy, runStudy, type StudyResult } from "humanish";
const parsed = parseStudy({
schema: STUDY_SCHEMA,
id: "library-example",
title: "Library example",
route: "computer-use",
subject: { source: "app-url", appUrl: "http://127.0.0.1:3000/" },
actor: { type: "openai-computer-use", mission: "Explore the app." },
execution: { target: "e2b-desktop", timeoutMs: 15_000 },
});
if (!parsed.ok) throw new Error(parsed.error.message);
const outcome = await runStudy(parsed.config, { cwd: process.cwd(), dryRun: true });
if (outcome.route === "computer-use") {
const result: StudyResult<"computer-use"> = outcome.result;
console.log(result.runId, result.ok);
}parseStudy takes a decoded object, not YAML text: a humanish.study.v3 study. It refuses a
0.107 humanish.lab.v2 one with HUMANISH_STUDY_V2_UNSUPPORTED; humanish migrate converts the
file. The StudyConfig it returns has the file's keys (route, actor, participants, caps,
mode) with defaults filled, so parsed.config.actor.type is the actor and
parsed.config.caps?.maxUsd the budget. A config you build yourself has the same keys.
runStudy refuses a config that still sets a StudyConfig field of humanish 0.110 (actors, a
scenario object, execution.caps, subject.topology, personas, or count, lanes, roster
or laneFocus on actor) with
HUMANISH_STUDY_V2_UNSUPPORTED, naming each field's v3 key. It refuses a config with no actor,
or a route its subject and actor do not take, with HUMANISH_STUDY_INVALID. Each refusal comes
in the route's result before anything runs. runStudy's options are RunStudyOptions:
env, scorer, createProvider, inProcess, prepareDesktop, onEvent, onStream,
analysisSignal and onObserverReady, beside dryRun, runId, count and rerun. The
library options contract
lists which routes accept each one and the refusals for the rest. routeOf(config) returns the
route a config takes, and StudyResult<route> is that route's result type.
onStream receives each participant's participantId, recordId (the id of its entry in
run.json simulations[], such as sim-001) and streamId, and on ready its sandboxId and
url.
Read a run
import { readFile } from "node:fs/promises";
import { renderObserver, verifyRun, type RunBundle } from "humanish";
const checked = await verifyRun(process.cwd(), "latest");
if (!checked.ok || !checked.bundlePath) throw new Error("the run did not verify");
const bundle = JSON.parse(await readFile(checked.bundlePath, "utf8")) as RunBundle;
console.log(bundle.review.verdict, checked.shareSafety.status);
await renderObserver(process.cwd(), checked.run, { open: false });bundle.review.verdict is what the participants experienced. A route result's ok, also
written to the run's status.json as outcome.ok, says whether the run worked as an
execution; outcome.execution.failures names each failure, such as a model provider's
cleanup that could not be confirmed. A run can read pass with ok: false.
This reads trusted local output. RunBundle["streams"][number] and
RunBundle["simulations"][number] name a bundle's members.
Bring a participant
A participant is a ComputerUseProvider (the brain) driving a ComputerUseExecutor (the desktop or app). Pass
createProvider to runStudy to replace the default brain on a computer-use study, and
inProcess: { executor } with it to drive an app in this process with no desktop. An
in-process run uses the same participant runner as hosted participants, so a declared
caps.maxUsd or caps.maxTotalUsd applies: report usage from your provider, or the
run stops at its first request with stopCause: "usage_unreported". The
participant example
runs one end to end:
npm install humanish
node node_modules/humanish/examples/participant/run.mjsrunComputerUseLoop is the loop underneath, for a caller that owns its own evidence;
defaultRedactionHooks is the redaction it requires. createOpenAiResponsesProvider builds the
default brain, and ComputerUseAdmissionLimitError is how a provider refuses a request before sending it.
Score a run
An AdapterScorerModule exports score, deriveFeedback and, on browser routes,
deriveArtifacts. Pass it as runStudy's scorer, or load the same file with
humanish run <study> --scorer <file>. The
scorer example does both.
Every scoring context carries ctx.bundle and ctx.runId.
A scorer written for one context passes through a helper that names it:
scorer: browserScorer({ score: (ctx: BrowserScoringContext) => … }) for computer use and
shared world, and terminalScorer(…) with TerminalProductScoringContext for terminal runs. An
inline scorer, or one that narrows ctx at runtime, passes as it is.
Both contexts name the study as ctx.studyId and the run as ctx.runId. A browser scorer's
context also carries ctx.route (computer-use or shared-world) and ctx.participantCount.
Removed exports
0.109.0 removed the 0.107 names that 0.108.0 kept as deprecated aliases: runLab,
parseLabConfig, LAB_CONFIG_SCHEMA, LabConfig, LabEvent, LabOutcome, LabResult,
LabRoute, RunLabOptions, BrowserLabScoringContext, the nine Cua* loop types and
CuaAdmissionLimitError. Import runStudy, parseStudy, StudyConfig, StudyEvent,
StudyOutcome, StudyResult, StudyRoute, RunStudyOptions, BrowserScoringContext, the
ComputerUse* types and ComputerUseAdmissionLimitError. It also removed labId on a study's
result (read studyId), simId on onStream events (read recordId), and
BrowserScoringContext.backend and laneCount (read route and participantCount). A scorer
reads ctx.studyId where it read ctx.labId, and a computer-use error's error.name is its
ComputerUse… class name.
The route runners runCuaActorLab, runScriptedBrowserLab, runTerminalProductLab and
runConcurrentSharedWorld, their option types, and the hook bag types are removed, along with the
cuaHooks, scriptedHooks, terminalHooks, sharedWorldHooks and automaticAnalysis options.
Use runStudy with the options above. A JavaScript caller that still passes a bag is refused with
HUMANISH_STUDY_OPTION_UNSUPPORTED, naming the option to use.
0.107.0 removed runDryRun, runCuaActorSession, the route result types, LabOutcome.backend,
LabBackend, selectLabBackend, the routesTo* predicates and the lab helpers that
parseStudy and routeOf cover. Use runStudy and its StudyResult<route>,
runComputerUseLoop, outcome.route and routeOf(config).
Library options
names the replacement for each.