Jev Harness / Jev Harness FAQ: installation, tools and verification

Jev Harness FAQ

简体中文 · Project overview · Task contract

These answers describe Apixly's Jev Harness v0.1.17. They distinguish the runtime's contracts, historical live evidence and illustrative media so readers can check each claim.

What is Jev Harness?

Jev Harness is a Python runtime for AI-authored task scripts with controlled tools and budgets. Goals may be open and routes unknown; new state supplies current candidates. The primary AI defines the task and registers tools; Jev chooses the next registered tool and its program-supplied parameter candidates. The program executes that bound action and independently checks whether the goal is complete. It supports semantic branching within explicit limits. See the project overview and architecture.

Does an open goal require a predefined route?

No. Availability and parameter providers can change after each observation; child choices can depend on selected parents. Bounds do not require knowing all future pages or arguments. The script defines capabilities rather than a task-specific next-step router. General evidence-reference reads and missing-argument exploration recovery remain priorities. New text needs an explicit trusted generator. See open goals and designed cases; their complete live acceptance is pending.

Is this an official TypeSafe AI product or the PyPI jev-harness package?

No. This is independent community software maintained by Apixly at apixly-ai/jev-harness. Its distribution name is apixly-jev-harness, its Python namespace is apixly_jev_harness, and its CLI aliases are apixly-jev-harness and jev-harness. The PyPI package named jev-harness is an unrelated project. Use the exact Git installation below. The architecture references identify other projects without implying endorsement or inherited benchmark results.

How does Jev Harness relate to Jev Filter?

Jev Filter provides typed semantic selection and filtering interfaces. Jev Harness builds a bounded observe–choose–execute–verify loop around registered tools, parameter domains and caller-owned completion checks. It pins Jev Filter v0.4.1 for its typed-batch inference transport and browser/desktop surfaces; it does not duplicate that inference implementation. Use a direct filtering interface when the work is selecting records; use the harness when a task needs multiple bounded tool decisions with independently observed progress. See architecture and the adoption sequence.

How do I install the exact v0.1.17 release?

Python 3.10+ and Git are required. Create a virtual environment and install the tagged repository URL:

python -m venv .venv
. .venv/bin/activate
pip install 'git+https://github.com/apixly-ai/jev-harness.git@v0.1.17'
jev-harness spec
jev-harness init-task task.py
jev-harness check-task task.py

Installation, spec, init-task and check-task make no model calls; downloading packages still needs network access. Running a task with Jev inference is billable and needs a caller-supplied credential through TYPESAFE_API_KEY or TYPESAFE_API_KEY_FILE. Use @main only for active development. Follow the first-task instructions before running.

Can Jev invent commands, selectors or tool arguments?

No. Jev may choose only registered tools and supplied parameter candidates. A candidate can be a static value or come from a trusted read-only provider derived from the current observation. Option(id, description, value) separates the model-visible description from the actual bound value; depends_on builds a child domain after its parent arguments are chosen. Singleton domains bind without inference. Model output does not become executable code, arbitrary text arguments, commands, paths, selectors or coordinates. Business tools still validate their own conditions. See the parameter protocol.

Provider arity uses the entry signature with follow_wrapped=False while respecting explicit signature. Generic *args forwarders with fewer than three explicit slots retain the advertised delegate contract, and selected two/three arguments must bind both. This preserves transparent two-argument forwarding while supporting explicit bridges; it does not change Task.tool signatures or prove every decorator's metadata/body behavior. For ambiguous wrappers use a clear two/three-argument entry or accurate Signature. See the provider evidence; no async, speed or savings claim.

Each complete parameter domain is deep-copied after duplicate-ID/size checks and before selection, so shared source mutation does not replace an ordinary Option/dict/list value while that choice waits. Providers remain fresh calls with their original identity, not a cache. It does not make acquisition atomic or prove global/business freshness, and custom deepcopy may fail even on an unselected value. Complete-domain copying costs additional time/memory based on all candidate values' object structure; failures and unknown usage remain. See the snapshot evidence, without a speed or fee promise.

Task.bind/execute also require a standard Action matching its latest offered snapshot under builtin typed comparison: bool differs from numbers, list/tuple differ, and finite 1/1.0 remains equivalent. This is not byte/object identity. Enum/str/dict/list subclasses may now be rejected even unchanged; normalize before offering, not after binding. Direct calls keep the existing ValueError codes; Engine execute rejection can remain uncertain even before the tool body. Keep prior effects/unknown usage and reconcile instead of retrying. Standard Engine does not actively mutate Actions. See the offered-action contract and offline evidence; no model-attack, general wrong-argument, semantic quality, speed or fee claim follows.

What tasks are suitable for the harness?

Use it for a bounded task with meaningful semantic choices, trusted tools and an observable completion condition: choosing a prerequisite tool, gathering evidence before a later action, or preparing a permitted synthetic browser draft. Keep exact arithmetic and fixed sequences in code. Keep open-ended writing in the primary AI or an explicitly registered trusted generation tool. The harness has no built-in general form order and candidate position does not establish priority; the caller supplies scope, preconditions and ordering. See next-step tool selection and the documented cases.

How does Jev Harness know a task is complete?

A done result requires the caller's independent verifier to pass against observed state. A model-only completion returns needs_review. Selecting a tool, submitting an action or receiving a tool result does not itself prove completion. A safe result summary admitted through result_context can advance the next decision without replacing verification. The terminal packet preserves blocked, missing-context, review and uncertain-execution outcomes. See context and results and terminal states.

verify_state uses recursive builtin matching: booleans differ from numbers, numeric 1/1.0 equivalence remains, nested list/tuple kinds stay distinct, and typed dictionary keys preserve the boolean/number distinction. Unsupported types/subclasses do not match; text checks require observed builtin strings. Expected text accepts None/builtin str; other types/subclasses fail construction, while unsupported comparison values do not match. This narrows implicit custom equality; use a custom verify for other deliberate semantics. Rejecting false completion can cause later tools and extra decisions, without changing authorization or proving speed, savings, AI accuracy or real UI acceptance. See the typed-verification evidence.

Is the homepage browser demo a live Jev recording?

The MP4 and GIF are labeled illustrative screenshot replays, made from current headless Camofox captures of the unchanged local synthetic form. They make no new model calls and are not continuous screen recordings. The separate historical live acceptance used Jev 1.13.0 with Camofox in v0.1.0, reached the final state in four tool executions, and independently checked the final DOM text and fields. Early review and authorization stops remain in the evidence. No real bookings or messages were sent. Read the media scope and acceptance ledger.

Are browser and native desktop tasks supported and verified?

The browser adapter uses a running loopback Camofox service, fresh observed targets, allowed origins and caller policy for exact click authorization. Native desktop adapters use pinned macOS Accessibility or Windows UI Automation backends with a named app/window; the adapter does not support Linux native desktop. Native desktop completion remains unverified: the Mac acceptance host was locked, and no live Windows acceptance is claimed. Offline adapter tests do not establish native completion. See adapters and evidence limits.

Does Jev Harness guarantee lower cost or faster completion?

No. Real Jev inference is billable. Reports retain whole-operation timing, returned context, known usage grouped by returned model where available, and unknown usage when evidence is missing. Missing model identity or interrupted attempts are not counted as zero cost. Published small development fixtures do not establish universal quality, latency or savings. Monetary estimates are not invoices and assume no current tariff. Check the measurement scope and negative results for each experiment.

Request counts are not invoices. Native Choice's existing telemetry.requests counts logical packed requests; optional owned_transport records this run's owned SDK-client attempts and HTTP client-object creation/reuse at terminal finish. Retries can make one logical request count as three attempts; bad credentials can leave it at zero attempts. Invalid counters are null, and absent metadata is unknown. Even complete counters do not prove HTTP delivery, server receipt, inference, billing, TCP connections or other tool traffic; unknown usage stays unknown. See the counter contract and A/B scope.

Is the runtime a sandbox, or does it guarantee exactly-once effects?

No. Task scripts are trusted Python; valid configuration still imports them and may cause top-level effects. Static check-task parses and compiles without executing a script, but passing it does not prove runtime validity. Within one run, the same observation, action ID and actual bound parameters are blocked before a repeated execution; this is an execution guard, not a cross-run exactly-once or resume guarantee. The caller's business tools own identity, authorization, idempotency, transactions and external reconciliation. Uncertain effects are not automatically retried. See static checking, execution context and terminals.

An uncertain DecisionStop result within the current execution window carries the available current bound action_id for private journal lookup and target reconciliation. It is not success/count proof, an idempotency key or permission to rerun. Bound arguments are not added to the terminal, and unknown usage remains unknown. See the identity evidence.

What execution and context limits should task authors enforce?

Inference has no result cache, and independent inference concurrency is capped at 30. The synchronous runtime uses cooperative wall-budget checkpoints; each tool must bound its own I/O and polling. At the top-of-cycle false-verification checkpoint, the existing step-limit check comes first, then an expired budget detected by this check stops before candidate acquisition. A late true verifier result still finishes done; after false verification, max_steps retains priority over this timeout check. Callbacks are not interrupted and not every boundary is checked, so this is not a hard deadline. Prior effects/unknown usage remain. Empty/raising candidate callbacks that are skipped no longer provide their former blocked/failed diagnostics; see the candidate-budget evidence, without a general speed or fee promise. Repeating an identical call against a fixed observation is unsupported. The default context policy keeps eight recent admitted action results plus an older action ledger; if admitted context still exceeds the conservative 15,000-character budget, it returns review rather than silently dropping constraints. Keep raw evidence in private journals and admit only safe summaries. See context management and the synchronous callback contract.

What should an AI do when a selection is not dispatched?

Read planning in the current selection's telemetry, and the loop diagnostic's tool or parameter stage. request_needs_narrowing indicates typed input needing narrowing; request_needs_context indicates required context/fields are missing. Unknown planning statuses use generic request_planning_review. Repair the program's candidates/descriptions or required fields while keeping the goal and safety constraints, rather than blindly retrying. The native runner was not entered for that selection, but earlier loop calls/effects or unknown usage may remain. Limits, prompts and model choices are unchanged; provider/HTTP413 failures and generic full-spec validation exceptions are separate. See the planning contract and offline evidence, without a semantic-quality, speed or fee claim.

View this page’s source on GitHub ↗