diff --git a/.env.example b/.env.example index 58acac1..dd5ef65 100644 --- a/.env.example +++ b/.env.example @@ -39,6 +39,13 @@ MODEL_NAME=google/gemini-2.5-flash # CUA_S1_SUBFOLDER=text # the text-only checkpoint; the window is 256 bytes of state # CUA_S1_DEVICE=auto # cpu, cuda, mps +# ---- OmniJev (the in-process vision decision model behind --model omnijev; needs `uv sync --extra omnijev`) ---- +# OMNIJEV_REPO=/path/to/OmniJev # required: a clone of https://github.com/tinnel123666888/OmniJev +# OMNIJEV_CHECKPOINT=tinnel123/OmniJev-0.8B # or OmniJev-2B, OmniJev (4B); a Hugging Face id or a local directory +# OMNIJEV_REVISION=v1.1 # for a Hugging Face id +# OMNIJEV_BASE=Qwen/Qwen3.5-0.8B # the Qwen3.5 backbone of the same size +# OMNIJEV_PROMPT=full # full: text state + goal + rules; short: goal + ask over the screenshot + # ---- Cua Driver (the hands of the desktop agents, e.g. calculator; install: https://cua.ai/docs/cua-driver) ---- # CUA_DRIVER_BIN=/usr/local/bin/cua-driver # when cua-driver is not on PATH # CUA_DRIVER_BIN=C:/path/to/cua-driver.exe # Windows: full path to the installed executable diff --git a/CHANGELOG.md b/CHANGELOG.md index 7b1eca6..3d9219a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,12 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); ver ### Added +- `--model omnijev` on the browser agents: [OmniJev](https://github.com/tinnel123666888/OmniJev) (Apache-2.0), a + Qwen3.5 vision-language decision model, in process behind `uv sync --extra omnijev` and a local clone named by + `OMNIJEV_REPO`. It decides over a screenshot; `docs/decision-models.md`, `docs/configuration.md`. +- Browser front: a decision model that reads images (`supports_images`) gets a viewport PNG with every tick's + observation, captured through the probe's run-code executor; the tick records its `screenshot_ms`. Jev, Laya and + Cua-S1 read text only and see no change. - The MCP `decide` tool accepts `model="jev"|"laya"|"cua"`, defaulting to `jev`. Local backends use their optional extras and need no Jev API key. - `docs/benchmarks.md`: the Google Flights driver comparison rerun on 2026-09-23 from Poland, every arm three times on diff --git a/README.md b/README.md index 1df99dd..4f846aa 100644 --- a/README.md +++ b/README.md @@ -56,6 +56,28 @@ wall clock. The other Allrecipes runs, longer games and the Google Flights drive +### OmniJev: a decision model that reads the screen + +[OmniJev](https://github.com/tinnel123666888/OmniJev) (Apache-2.0, Beijing Zhongguancun Academy, CASIA and Zevo) is a +System 1 decision model on Qwen3.5 vision-language backbones (0.8B, 2B, 4B): it answers the same typed questions as +Jev over a screenshot, a video or a robot camera. `--model omnijev` puts it in the slot of the browser agents, which +then send it a screenshot of the page at every step ([docs/decision-models.md](docs/decision-models.md)). + +The clips below are OmniJev's own v1.1 demos with its 4B model: replays of recorded trajectories with the model's +probabilities, not s1a runs and not live control. They are shown from the OmniJev repository and keep their upstream +terms. + + + + + + + + + + +
OmniJev on Mind2Web web tasks: a Central Park to JFK route and a nightstand comparison, with the model's probabilities per step
Web: Mind2Web test tasks (OmniJev v1.1 replay)
OmniJev on AndroidControl phone tasks: London weather, a 59-minute timer and a drawing tutorial
Phone: AndroidControl tasks (OmniJev v1.1 replay)
OmniJev on Atari recordings: Enduro, Skiing and Pong, next input and steering
Games: Enduro, Skiing and Pong (OmniJev v1.1 replay)
OmniJev on a two-view Bridge robot recording folding a cloth: jog direction, gripper and move size
Robotics: folding a cloth, two camera views (OmniJev v1.1 replay)
+ ## Choose your path ### From Claude Code or Codex @@ -124,8 +146,8 @@ The gates and the templates: [docs/skills.md](docs/skills.md#build-a-system-1-ag - `game2048`, `millionaire`, `blackjack`: games with a score per episode. - `injection_guard`: a rail that answers one question at a hook of a running agent and fails closed. -Every agent runs on `jev`, `laya` or `cua`, and on the chat model for the comparison. Flags, run commands and -extras: [docs/agents.md](docs/agents.md). +Every agent runs on `jev`, `laya` or `cua`, and on the chat model for the comparison; the browser agents also run on +`omnijev`, which decides over a screenshot. Flags, run commands and extras: [docs/agents.md](docs/agents.md). ## How it works diff --git a/docs/agents.md b/docs/agents.md index c7ff39a..9cd9033 100644 --- a/docs/agents.md +++ b/docs/agents.md @@ -18,7 +18,7 @@ hook of a running agent. The injection guard rail fails closed: a decision error Every tool agent takes `--model jev|laya|cua|llm|random|rule`, `--rethink on|off`, `--episodes N`, `--seed S`, `--max-steps` and `--timeout`, and writes a Harbor-shaped job folder under `evals/results//`. A browser agent -takes `--model jev|laya|cua|llm` and `--goal`. A rail takes `--model jev|laya`, the two models that answer `noul`. +takes `--model jev|laya|cua|omnijev|llm` and `--goal`. A rail takes `--model jev|laya`, the two models that answer `noul`. `uv run python -m evals.table evals/results` aggregates every job folder per eval and model into one table. Every `run` prints one JSON object on stdout and nothing else there; `s1a-mcp` serves the same agents over stdio diff --git a/docs/architecture.md b/docs/architecture.md index 5bfb690..3e9125a 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -62,7 +62,7 @@ the first import of both, routes the harness logs to files under `runs/logs` bef MCP stdio protocol only. Every `run` prints one JSON object on stdout: a tool agent's series summary with its `job_dir`, a browser agent's -answer, a rail's evaluation. A browser agent takes `--model jev|laya|cua|llm`; its policy switches are run-time flags: +answer, a rail's evaluation. A browser agent takes `--model jev|laya|cua|omnijev|llm`; its policy switches are run-time flags: `--batch on|off`, `--prefetch on|off`, `--goal-values on|off`. `s1a-mcp` serves the same agents to an MCP host over stdio, one Runner for the server's lifetime and one run at a time. `uv run python -m evals.table evals/results` aggregates every job folder per eval and model into one table. `scripts/showcase.sh` plays one visual episode per @@ -71,7 +71,7 @@ eval and model outside the matrix and `python -m evals.replay` renders a pair si Every tool agent, `desktop` included, takes `--model jev|laya|cua|llm|random|rule`, `--rethink on|off`, `--episodes N`, `--seed S`, `--max-steps`, `--timeout` and `--headed`, and writes a Harbor-shaped job folder under -`evals/results//`. Every browser agent takes `--model jev|laya|cua|llm` and `--goal`. A rail takes +`evals/results//`. Every browser agent takes `--model jev|laya|cua|omnijev|llm` and `--goal`. A rail takes `--model jev|laya`, the two models that answer `noul`. `decide` and `probe` take `--model jev|laya|cua`. On a browser agent `laya` needs `LAYA_MAX_LEN` raised to the page's size; `cua` reads a 256-byte context (header, goal, state, then rules) and 96 bytes per option, a baseline on any page. Exit codes: 0 for a finished run, including one whose diff --git a/docs/configuration.md b/docs/configuration.md index 7c71622..a6b575b 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -35,6 +35,11 @@ Variables can be exported in your shell or placed in a `.env` file at the root o | `CUA_S1_CHECKPOINT` | `cua` model | `cua-ai/cua-s1-nano-0.1` | Hugging Face checkpoint ID or local directory for Cua-S1 Nano option scorer. | | `CUA_S1_SUBFOLDER` | `cua` model | `text` | Subfolder within checkpoint directory containing text option scoring weights. | | `CUA_S1_DEVICE` | `cua` model | `auto` | PyTorch device used for Cua-S1 Nano evaluation (`auto`, `cpu`, `cuda`, or `mps`). | +| `OMNIJEV_REPO` | `omnijev` model | *(unset, required)* | Local clone of the OmniJev repository; its `mso` package is imported from there. | +| `OMNIJEV_CHECKPOINT` | `omnijev` model | `tinnel123/OmniJev-0.8B` | Hugging Face ID or local directory of the OmniJev adapter and heads (`OmniJev-2B`, `OmniJev` for 4B). | +| `OMNIJEV_REVISION` | `omnijev` model | `v1.1` | Revision of a Hugging Face `OMNIJEV_CHECKPOINT`. | +| `OMNIJEV_BASE` | `omnijev` model | `Qwen/Qwen3.5-0.8B` | Hugging Face ID or local directory of the Qwen3.5 backbone of the same size as the checkpoint. | +| `OMNIJEV_PROMPT` | `omnijev` model | `full` | `full` sends the text state, goal and rules with each question; `short` sends the goal and the ask only, the screenshot carrying the page. | | `CUA_DRIVER_BIN` | `desktop` agent | `cua-driver` | Path to the `cua-driver` executable on Windows or macOS when not located on `PATH`. | | `CUA_DRIVER_PERMISSION_MODE` | `desktop` agent | `standard` | Permission mode passed to `cua-driver mcp` (`standard`, or `bounded` for restricted capability manifests). | | `HF_HOME` | Hugging Face runtime | `~/.cache/huggingface` | Cache directory where Laya and Cua-S1 checkpoints are downloaded on first run. | diff --git a/docs/decision-models.md b/docs/decision-models.md index 89b2e0d..b200c84 100644 --- a/docs/decision-models.md +++ b/docs/decision-models.md @@ -45,6 +45,7 @@ shorthands; `warm()` and `close()` open and release the backend. | `jev` | `JevModel(transport)` | `jev` | the request body every front sent before the layer existed, byte for byte; `from_env` picks TypeSafe or the OpenRouter proxy (see [configuration.md](configuration.md)) | | `laya` | `LayaModel(agent, model=)` | `laya` | one forward pass per call on a thread; `MODEL_SERVICE_CONFIG_ERROR` when `input_tokens` fills the window (Laya cuts the state silently; see `LAYA_MAX_LEN` and `LAYA_HEAD_MAX_LEN` in [configuration.md](configuration.md)); `ValueError` and `RuntimeError` from the library become `MODEL_CALL_FAILED` | | `cua` | `CuaS1Model(scorer, collator, model=, context_bytes=, option_bytes=)` | `cua` | Cua-S1 Nano, one `score_elements` pass per request on a thread; choice questions only, text only, deterministic; the context is header, state and rules; the checkpoint reads its first 256 bytes, and the first overflowing request logs one warning; `from_env` reads `CUA_S1_*` (see [configuration.md](configuration.md)) | +| `omnijev` | `OmniJevModel(agent, model=)` | `omnijev` | OmniJev (a Qwen3.5 vision-language decision model), one `MSO1.system_one` call per request on a thread; choice and noul questions over the observation's screenshot, deterministic; the text state, goal and rules go in front of each question (a browser state puts its recent actions first and caps the page text at 2,000 characters, so a long page cannot push the history out of the 6,000-character context), a browser element row becomes its label and value; a choice's probabilities, which OmniJev leaves `abstain` out of, are renormalised over the offered keys; an observation without an image is `MODEL_SERVICE_CONFIG_ERROR`; `from_env` reads `OMNIJEV_*` (see [configuration.md](configuration.md)) | | `random` | `RandomModel(seed)` | `random` | uniform over the offered keys, confidence 0, one seeded stream per episode; choice questions only | | `rule` | `RuleModel(name, rule)` | the rule's name | one-hot, confidence 1; a key outside the menu raises `RuntimeError`, a bug in the rule | diff --git a/pyproject.toml b/pyproject.toml index 682ddda..2c1d713 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -45,7 +45,17 @@ cua = [ # Cua-S1 Nano behind --model cua; pinned to the Cua PR that ships the c "cua-s1 @ git+https://github.com/trycua/cua.git@aea61b6eb97e2d8c0f6f71eb804e5769fe910af4#subdirectory=libs/cua-s1/python", "huggingface-hub>=0.24", ] -dev = ["pytest>=8", "pytest-asyncio>=0.24", "ruff>=0.6", "ty>=0.0.83"] +omnijev = [ # OmniJev behind --model omnijev; the model code itself is a clone named by OMNIJEV_REPO (not a package) + "torch>=2.4", + "torchvision>=0.15", # transformers' Qwen3-VL video processor imports it; OmniJev's requirements.txt omits it + "transformers>=5.0", + "peft>=0.15", + "accelerate", + "safetensors", + "pillow", + "huggingface-hub>=0.24", +] +dev =["pytest>=8", "pytest-asyncio>=0.24", "ruff>=0.6", "ty>=0.0.83"] [build-system] requires = ["hatchling"] @@ -76,6 +86,7 @@ include = [ "s1a/agents/blackjack.py", "s1a/decision_models/cua.py", "s1a/decision_models/laya.py", + "s1a/decision_models/omnijev.py", ] [tool.ty.overrides.rules] diff --git a/s1a/browser/browse.py b/s1a/browser/browse.py index b9be304..3eb1fac 100644 --- a/s1a/browser/browse.py +++ b/s1a/browser/browse.py @@ -38,8 +38,9 @@ "jev", "laya", "cua", + "omnijev", "llm", -) # a decision model (Jev over HTTP, Laya or Cua-S1 in process) or the chat model +) # a decision model (Jev over HTTP, Laya, Cua-S1 or OmniJev in process) or the chat model RUNS_DIR = HOME / "runs" / "browser" @@ -164,7 +165,7 @@ async def browse( workspace = str(logs_dir / "workspace") # the harness scaffolds SOUL.md, memory/ and friends here, not in the cwd instance = BrowserInstanceConfig(launch_args=browser_launch_args(headless)) match model_name: - case "jev" | "laya" | "cua": + case "jev" | "laya" | "cua" | "omnijev": if decision_model is None: raise RuntimeError(f"--model {model_name} needs a decision model") slot_model = BrowserDecisionModel(spec, policy, counted, decision_model=decision_model, value_model=None) @@ -224,7 +225,7 @@ def parser(spec: BrowserAgentSpec) -> argparse.ArgumentParser: "--model", choices=BROWSER_MODEL_NAMES, required=True, - help="who decides each browser step: jev (over HTTP), laya or cua (in process), or llm (the chat model in MODEL_NAME)", + help="who decides each browser step: jev (over HTTP), laya, cua or omnijev (in process; omnijev reads a screenshot), or llm (the chat model in MODEL_NAME)", ) build.add_argument( "--goal", default=spec.goal, required=spec.goal is None, help="the task; the spec's goal when it has one" diff --git a/s1a/browser/decision_model.py b/s1a/browser/decision_model.py index 001478b..e983647 100644 --- a/s1a/browser/decision_model.py +++ b/s1a/browser/decision_model.py @@ -15,6 +15,7 @@ from __future__ import annotations import asyncio +import base64 import json import re import statistics @@ -41,7 +42,7 @@ top_probabilities, ) from s1a.browser.probe_js import POLICY_PROBE_JS, STAMP_ATTRIBUTE -from s1a.decision_models import DecisionModel +from s1a.decision_models import DecisionModel, Image, Observation from s1a.spec import BrowserAgentSpec BROWSER_TURN_TOOL = "browser_click" @@ -82,6 +83,14 @@ MAX_PROBE_SETTLE_MS = 1500 # keeps load(3s)+settle+1s JS lastResort >=1s under the 30s transport request timeout WAIT_SETTLE_BUDGET_MS = 3000 # total in-page settle time one WAIT streak may spend before the step gives up URL_RE = re.compile(r"https?://[^\s'\"<>]+") +# The viewport as PNG, for a decision model that reads images; one run-code call through the probe's own executor. +# A background tab renders no frames, so the page is brought to the front first; the capture waits 15 s at most. +SCREENSHOT_JS = ( + "async (page) => { await page.bringToFront(); " + 'const png = await page.screenshot({type: "png", timeout: 15000}); ' + "return JSON.stringify({png: png.toString('base64')}); }" +) +_SCREENSHOT_RE = re.compile(r'png\\*"\s*:\s*\\*"([A-Za-z0-9+/=]+)') @dataclass(frozen=True) @@ -258,6 +267,11 @@ async def _decide_message(self, messages: Any) -> AssistantMessage: run.values = [] values = run.values observation = build_observation(space, snapshot, run.history) + screenshot_ms = None + if self._decision_model.supports_images: + screenshot, screenshot_ms = await self._screenshot() + if screenshot is not None: + observation = Observation(observation.state, images=(screenshot,)) questions = build_questions( space, goal=run.goal, values=values, rules=self._spec.rules, language=self._language ) @@ -287,6 +301,8 @@ async def _decide_message(self, messages: Any) -> AssistantMessage: "probabilities": probabilities, # the replay's bars: the settled head's top keys "candidates": candidates, } + if screenshot_ms is not None: # an image-reading model: the capture is part of the step's cost + record["screenshot_ms"] = screenshot_ms run.ticks.append(record) if move.operation == "WAIT": run.consecutive_waits += 1 @@ -397,6 +413,24 @@ async def _finish_probe( self._prefetch_values(run, snapshot) return snapshot, round((time.perf_counter() - started) * 1000) + async def _screenshot(self) -> tuple[Image | None, int]: + """The viewport as PNG and the milliseconds the capture took. ``None`` when it fails: the step goes on + without the picture and the decision model says whether it can decide without one.""" + started = time.perf_counter() + executor = getattr(self._runtime, "code_executor", None) + raw: Any = None + if callable(executor): + try: + raw = await executor(SCREENSHOT_JS) + except Exception: # noqa: BLE001 - a failed capture degrades to a text-only observation + logger.warning("[BrowserDecisionModel] screenshot failed", exc_info=True) + ms = round((time.perf_counter() - started) * 1000) + found = _SCREENSHOT_RE.search(json.dumps(raw, default=str)) if raw is not None else None + if found is None: + logger.warning("[BrowserDecisionModel] no screenshot in the run-code result") + return None, ms + return Image(base64.b64decode(found.group(1))), ms + async def _raw_probe(self, *, settle_ms: int, quiet_ms: int, after: dict[str, Any] | None) -> dict[str, Any]: params = { "stamp_attribute": STAMP_ATTRIBUTE, diff --git a/s1a/decision_models/__init__.py b/s1a/decision_models/__init__.py index bf68c9f..0142db5 100644 --- a/s1a/decision_models/__init__.py +++ b/s1a/decision_models/__init__.py @@ -14,6 +14,7 @@ from s1a.decision_models.fakes import ScriptedModel, ScriptedTransport from s1a.decision_models.jev import JevModel, jev_question from s1a.decision_models.laya import LayaModel +from s1a.decision_models.omnijev import OmniJevModel from s1a.decision_models.types import ( Answer, Choice, @@ -47,6 +48,7 @@ "LayaModel", "Noul", "NoulQuestion", + "OmniJevModel", "Observation", "Question", "RandomModel", diff --git a/s1a/decision_models/factory.py b/s1a/decision_models/factory.py index d81ae45..50a1a61 100644 --- a/s1a/decision_models/factory.py +++ b/s1a/decision_models/factory.py @@ -8,18 +8,20 @@ from s1a.decision_models.cua import CuaS1Model from s1a.decision_models.jev import JevModel from s1a.decision_models.laya import LayaModel +from s1a.decision_models.omnijev import OmniJevModel DECISION_MODEL_NAMES = ( "jev", "laya", "cua", + "omnijev", "random", "rule", ) # the names that build a decision model; ``llm`` is not one def build_model(model_name: str, *, seed: int = 0, rule: tuple[str, Rule] | None = None) -> DecisionModel: - """``jev``, ``laya`` and ``cua`` from the environment, ``random`` from the seed, ``rule`` from the agent's baseline.""" + """``jev``, ``laya``, ``cua`` and ``omnijev`` from the environment, ``random`` from the seed, ``rule`` from the agent's baseline.""" match model_name: case "jev": return JevModel.from_env() @@ -27,6 +29,8 @@ def build_model(model_name: str, *, seed: int = 0, rule: tuple[str, Rule] | None return LayaModel.from_env() # the laya import happens inside case "cua": return CuaS1Model.from_env() # the cua_s1 import happens inside + case "omnijev": + return OmniJevModel.from_env() # the OmniJev import happens inside; it decides over a screenshot case "random": return RandomModel(seed) case "rule": diff --git a/s1a/decision_models/omnijev.py b/s1a/decision_models/omnijev.py new file mode 100644 index 0000000..64f588f --- /dev/null +++ b/s1a/decision_models/omnijev.py @@ -0,0 +1,200 @@ +# coding: utf-8 +"""OmniJev in process: a Qwen3.5 vision-language decision model that reads a screenshot, one call per request on a thread. + +OmniJev (https://github.com/tinnel123666888/OmniJev, Apache-2.0) answers the TypeSafe question shape over an image: +``MSO1.system_one({"images": [path]}, questions)``. It ships as a repository, not a package, so ``OMNIJEV_REPO`` +names a local clone; torch, transformers and peft come from the ``omnijev`` extra. Everything heavy is imported inside +``from_env``; the module imports without the extra. +""" + +from __future__ import annotations + +import asyncio +import json +import os +import sys +import tempfile +import time +from pathlib import Path +from typing import Protocol + +from openjiuwen.core.common.exception.codes import StatusCode +from openjiuwen.core.common.exception.errors import build_error + +from s1a.decision_models.base import DecisionModel +from s1a.decision_models.types import ChoiceQuestion, Image, Json, Observation, Question, Reply + +OMNIJEV_DEFAULT_CHECKPOINT = "tinnel123/OmniJev-0.8B" +OMNIJEV_DEFAULT_REVISION = "v1.1" +OMNIJEV_DEFAULT_BASE = "Qwen/Qwen3.5-0.8B" +OMNIJEV_STATE_CHARS = 6000 # the text state read with the picture; the picture carries the rest +OMNIJEV_PAGE_TEXT_CHARS = 2000 # a browser page's free text, capped on its own so the actions and elements stay in +OMNIJEV_PROMPTS = ("full", "short") # full: text state + goal + rules + ask; short: goal + ask over the screenshot +OMNIJEV_DEFAULT_PROMPT = "full" +_SUFFIXES = {"image/png": ".png", "image/jpeg": ".jpg", "image/webp": ".webp"} + + +class OmniJevAgent(Protocol): + """The slice of ``mso.infer.MSO1`` the adapter uses.""" + + def system_one(self, state: Json, questions: dict[str, Json]) -> dict[str, Json]: ... + + +def omnijev_option(description: str | Json) -> str: + """One option's description as text: a browser element row (``{"element": "[12] Where from?", ...}``) becomes its + label and value; any other dict becomes compact JSON.""" + if isinstance(description, str): + return description + if "element" in description: + label = str(description["element"]).split("] ", 1)[-1] + if description.get("option"): + label += f" / {description['option']}" + value = str(description.get("current_value") or "") + return label + (f" = {value}" if value else "") + return json.dumps(description, ensure_ascii=False) + + +def omnijev_context(observation: Observation) -> str: + """The observation's text state, read in front of every question (OmniJev reads text context in the instructions). + + A browser state is reordered so the whole budget cannot go to the page's free text: the recent actions come first + (the screenshot cannot show them, and the rules depend on them, e.g. PRESS_ENTER after a Search click that did + nothing), then the page with its text capped at ``OMNIJEV_PAGE_TEXT_CHARS``, then the element table.""" + state = observation.state + if isinstance(state, dict) and isinstance(state.get("page"), dict): + page = dict(state["page"]) + page_text = str(page.get("text") or "") + if len(page_text) > OMNIJEV_PAGE_TEXT_CHARS: + page["text"] = page_text[:OMNIJEV_PAGE_TEXT_CHARS] + " …" + rest = {key: value for key, value in state.items() if key not in ("recent_actions", "page")} + state = {"recent_actions": state.get("recent_actions", []), "page": page, **rest} + text = state if isinstance(state, str) else json.dumps(state, ensure_ascii=False, separators=(",", ":")) + return text[:OMNIJEV_STATE_CHARS] + + +def omnijev_question(question: Question, context: str, *, prompt: str = "full") -> Json: + """One question in OmniJev's shape. ``full``: the text state, the goal and the rules in front of the ask. + ``short``: the goal and the ask only; the screenshot carries the page and the rules stay out.""" + full = prompt == "full" + lines = [f"State: {context}"] if context and full else [] + if isinstance(question, ChoiceQuestion): + if question.goal: + lines.append(f"Task: {question.goal}") + if full: + lines.extend(question.rules) + lines.append( + f"Which element should {question.operation} act on?" if question.operation else "Which option comes next?" + ) + criteria = {key: omnijev_option(description) for key, description in question.options.items()} + return {"type": "choice", "instructions": "\n".join(lines), "criteria": criteria} + lines.append(question.question) + for verdict, text in (question.criteria or {}).items(): + lines.append(f"{verdict}: {text}") + return {"type": "noul", "instructions": "\n".join(lines)} + + +def omnijev_answer(answer: Json) -> Json: + """OmniJev's answer in the shape ``decide_many`` validates. A choice's probabilities sum to ``1 - abstain`` (its + "none of these"); they are renormalised over the keys it returned and the confidence recomputed with Jev's + definition on them. The key stays OmniJev's own, so a malformed answer still fails validation. ``abstain`` stays + in the reply's raw payload.""" + if "noul" in answer: + return {"noul": float(answer["noul"])} + raw = {key: float(p) for key, p in (answer.get("probabilities") or {}).items()} + total = sum(raw.values()) + if not raw or total <= 0: + return {"choice": answer.get("choice"), "probabilities": raw, "confidence": 0.0} + probabilities = {key: p / total for key, p in raw.items()} + k = len(probabilities) + peak = max(probabilities.values()) + confidence = 1.0 if k < 2 else max(0.0, min(1.0, (k * peak - 1) / (k - 1))) + return {"choice": answer.get("choice"), "probabilities": probabilities, "confidence": confidence} + + +class OmniJevModel(DecisionModel): + """OmniJev's ``MSO1`` behind the interface: choice and noul questions over one screenshot, deterministic.""" + + name = "omnijev" + supports_images = True + deterministic = True + + def __init__(self, agent: OmniJevAgent, *, model: str, prompt: str = OMNIJEV_DEFAULT_PROMPT) -> None: + if prompt not in OMNIJEV_PROMPTS: + raise ValueError(f"omnijev prompt is one of {OMNIJEV_PROMPTS}, not {prompt!r}") + self._agent = agent + self._model = model + self._prompt = prompt + + @property + def model(self) -> str: + return self._model + + async def _decide(self, observation: Observation, questions: dict[str, Question]) -> Reply: + if not observation.images: + raise build_error( + StatusCode.MODEL_SERVICE_CONFIG_ERROR, + error_msg="omnijev decides over a screenshot; this observation carries no image", + ) + context = omnijev_context(observation) + asked = {name: omnijev_question(question, context, prompt=self._prompt) for name, question in questions.items()} + started = time.perf_counter() + with tempfile.TemporaryDirectory(prefix="s1a-omnijev-") as folder: + paths = [_write_image(image, Path(folder), index) for index, image in enumerate(observation.images)] + try: + payload = await asyncio.to_thread(self._agent.system_one, {"images": paths}, asked) + except (ValueError, RuntimeError, OSError) as exc: # a bad image, and torch failures + raise build_error( + StatusCode.MODEL_CALL_FAILED, cause=exc, error_msg=f"omnijev forward pass failed: {exc}" + ) from exc + ms = round((time.perf_counter() - started) * 1000) + answers = {name: omnijev_answer(payload[name]) for name in questions if name in payload} + return Reply(answers=answers, latency_ms=ms, model=self._model, raw=payload) + + @classmethod + def from_env(cls) -> "OmniJevModel": + """``OMNIJEV_REPO`` (a local clone of the OmniJev repository, required), ``OMNIJEV_CHECKPOINT`` (a hub id or a + local directory), ``OMNIJEV_REVISION`` (for a hub id) and ``OMNIJEV_BASE`` (the Qwen3.5 backbone of the same + size, a hub id or a local directory), ``OMNIJEV_PROMPT`` (``full`` or ``short``, see ``omnijev_question``).""" + prompt = os.getenv("OMNIJEV_PROMPT") or OMNIJEV_DEFAULT_PROMPT + if prompt not in OMNIJEV_PROMPTS: + raise build_error( + StatusCode.MODEL_SERVICE_CONFIG_ERROR, + error_msg=f"OMNIJEV_PROMPT is one of {', '.join(OMNIJEV_PROMPTS)}, not {prompt!r}", + ) + repo = os.getenv("OMNIJEV_REPO") + if not repo or not (Path(repo) / "mso" / "infer.py").is_file(): + raise build_error( + StatusCode.MODEL_SERVICE_CONFIG_ERROR, + error_msg=( + "--model omnijev needs OMNIJEV_REPO, a local clone of " + "https://github.com/tinnel123666888/OmniJev, and the omnijev extra: uv sync --extra omnijev" + ), + ) + sys.path.insert(0, str(Path(repo).resolve())) + try: + from mso.infer import MSO1 + except ImportError as exc: + raise build_error( + StatusCode.MODEL_SERVICE_CONFIG_ERROR, + error_msg=f"--model omnijev could not import OmniJev ({exc}); uv sync --extra omnijev", + ) from exc + checkpoint = os.getenv("OMNIJEV_CHECKPOINT") or OMNIJEV_DEFAULT_CHECKPOINT + revision = os.getenv("OMNIJEV_REVISION") or OMNIJEV_DEFAULT_REVISION + base = os.getenv("OMNIJEV_BASE") or OMNIJEV_DEFAULT_BASE + agent = MSO1(_local(checkpoint, revision), _local(base, None)) + return cls(agent, model=checkpoint if Path(checkpoint).is_dir() else f"{checkpoint}@{revision}", prompt=prompt) + + +def _local(name: str, revision: str | None) -> str: + """A local directory as it is; a hub id fetched once into the Hugging Face cache.""" + if Path(name).expanduser().is_dir(): + return str(Path(name).expanduser()) + from huggingface_hub import snapshot_download + + return snapshot_download(name, revision=revision) + + +def _write_image(image: Image, folder: Path, index: int) -> str: + path = folder / f"{index}{_SUFFIXES.get(image.media_type, '.png')}" + path.write_bytes(image.data) + return str(path) diff --git a/tests/test_browse.py b/tests/test_browse.py index 71f9c20..7504a5a 100644 --- a/tests/test_browse.py +++ b/tests/test_browse.py @@ -56,7 +56,7 @@ def test_a_timeout_or_max_steps_at_or_below_zero_is_a_usage_error(self) -> None: self.assertEqual(caught.exception.code, 2) def test_the_model_flag_takes_a_decision_model_or_the_chat_model(self) -> None: - self.assertEqual(browse.BROWSER_MODEL_NAMES, ("jev", "laya", "cua", "llm")) + self.assertEqual(browse.BROWSER_MODEL_NAMES, ("jev", "laya", "cua", "omnijev", "llm")) for model_name in browse.BROWSER_MODEL_NAMES: self.assertEqual(browse.parser(SPEC).parse_args(["--model", model_name, "--goal", "x"]).model, model_name) with self.assertRaises(SystemExit): diff --git a/tests/test_browser_screenshot.py b/tests/test_browser_screenshot.py new file mode 100644 index 0000000..c9f0b79 --- /dev/null +++ b/tests/test_browser_screenshot.py @@ -0,0 +1,49 @@ +# coding: utf-8 +"""``BrowserDecisionModel._screenshot``: the viewport PNG an image-reading decision model decides over.""" + +from __future__ import annotations + +import base64 +import json +from types import SimpleNamespace +from typing import Any +from unittest import IsolatedAsyncioTestCase + +from s1a.browser.decision_model import SCREENSHOT_JS, BrowserDecisionModel + +PNG = b"\x89PNG\r\n\x1a\n fake" +ENCODED = base64.b64encode(PNG).decode() + + +def _runtime(result: Any = None, error: Exception | None = None) -> SimpleNamespace: + calls: list[str] = [] + + async def executor(code: str) -> Any: + calls.append(code) + if error is not None: + raise error + return result + + return SimpleNamespace(code_executor=executor, calls=calls) + + +async def _shot(runtime: Any) -> tuple[Any, int]: + return await BrowserDecisionModel._screenshot(SimpleNamespace(_runtime=runtime)) # type: ignore[arg-type] + + +class TestScreenshot(IsolatedAsyncioTestCase): + async def test_the_png_comes_back_from_the_run_code_payload(self) -> None: + payload = {"__browser_compact_rpc__": True, "payload": {"result": json.dumps({"png": ENCODED})}} + runtime = _runtime(payload) + image, ms = await _shot(runtime) + self.assertEqual(runtime.calls, [SCREENSHOT_JS]) + self.assertEqual((image.data, image.media_type), (PNG, "image/png")) + self.assertGreaterEqual(ms, 0) + + async def test_a_failed_capture_is_no_image_not_an_error(self) -> None: + image, _ms = await _shot(_runtime(error=RuntimeError("page closed"))) + self.assertIsNone(image) + + async def test_a_result_without_a_png_is_no_image(self) -> None: + image, _ms = await _shot(_runtime({"payload": {"result": "{}"}})) + self.assertIsNone(image) diff --git a/tests/test_decision_models_factory.py b/tests/test_decision_models_factory.py index 3c85f6d..6673273 100644 --- a/tests/test_decision_models_factory.py +++ b/tests/test_decision_models_factory.py @@ -40,7 +40,7 @@ def test_every_name_builds_its_class(self) -> None: rule = build_model("rule", rule=("always-inc", lambda state, options: "inc")) self.assertIsInstance(rule, RuleModel) self.assertEqual(rule.name, "always-inc") - self.assertEqual(DECISION_MODEL_NAMES, ("jev", "laya", "cua", "random", "rule")) + self.assertEqual(DECISION_MODEL_NAMES, ("jev", "laya", "cua", "omnijev", "random", "rule")) def test_the_errors(self) -> None: with self.assertRaises(RuntimeError): diff --git a/tests/test_decision_models_omnijev.py b/tests/test_decision_models_omnijev.py new file mode 100644 index 0000000..a074323 --- /dev/null +++ b/tests/test_decision_models_omnijev.py @@ -0,0 +1,185 @@ +# coding: utf-8 +"""``OmniJevModel`` over a fake ``mso.infer.MSO1`` (no torch): the contract, the question and answer mapping, the +screenshot hand-off, the errors, and ``from_env`` without a clone.""" + +from __future__ import annotations + +import os +from pathlib import Path +from typing import Any +from unittest import IsolatedAsyncioTestCase, TestCase +from unittest.mock import patch + +from openjiuwen.core.common.exception.codes import StatusCode +from openjiuwen.core.common.exception.errors import BaseError + +import decision_model_contract as contract +from s1a.decision_models import ChoiceQuestion, DecisionModel, Image, Observation, OmniJevModel +from s1a.decision_models import omnijev as omnijev_module + +SCREEN = Image(b"\x89PNG fake screenshot") +PICTURED = Observation(contract.OBSERVATION.state, images=(SCREEN,)) + + +class FakeMSO1: + """``MSO1`` without torch: ``system_one`` in OmniJev's shapes (a choice leaves ``abstain`` out of its sum).""" + + def __init__(self, *, answers: list[dict[str, Any]] | None = None, error: Exception | None = None) -> None: + self.calls: list[tuple[dict[str, Any], dict[str, Any], list[bytes]]] = [] + self._answers = list(answers or []) + self._error = error + + def system_one(self, state: dict[str, Any], questions: dict[str, Any]) -> dict[str, Any]: + images = [Path(path).read_bytes() for path in state["images"]] # the files exist during the call + self.calls.append((state, questions, images)) + if self._error is not None: + raise self._error + if self._answers: + return self._answers.pop(0) + return {name: self._answer(question) for name, question in questions.items()} + + @staticmethod + def _answer(question: dict[str, Any]) -> dict[str, Any]: + if question["type"] == "noul": + return {"noul": 0.3, "latency_s": 0.01} + keys = list(question["criteria"]) + probabilities = {key: (0.6 if key == keys[-1] else 0.3 / max(1, len(keys) - 1)) for key in keys} + return {"choice": keys[-1], "probabilities": probabilities, "abstain": 0.1, "valid": True, "confidence": 0.5} + + +def _model(agent: FakeMSO1 | None = None) -> OmniJevModel: + return OmniJevModel(agent or FakeMSO1(), model="tinnel123/OmniJev-0.8B@v1.1") + + +class TestOmniJevContract(contract.DecisionModelContract, IsolatedAsyncioTestCase): + """The shared contract, over an observation that carries a screenshot: OmniJev decides over one.""" + + def setUp(self) -> None: + patcher = patch.object(contract, "OBSERVATION", PICTURED) + patcher.start() + self.addCleanup(patcher.stop) + + def make(self) -> DecisionModel: + return _model() + + def make_scripted(self, answers: list[dict[str, Any]]) -> DecisionModel: + return _model(FakeMSO1(answers=answers)) + + +class TestMapping(IsolatedAsyncioTestCase): + async def test_the_screenshot_reaches_omnijev_as_a_file_and_is_removed_after(self) -> None: + agent = FakeMSO1() + await _model(agent).decide_many(PICTURED, {"pick": contract.PICK}) + ((state, _questions, images),) = agent.calls + self.assertEqual(images, [SCREEN.data]) + self.assertFalse(any(Path(path).exists() for path in state["images"])) + + async def test_the_text_state_goal_and_rules_come_before_the_ask(self) -> None: + agent = FakeMSO1() + question = ChoiceQuestion({"1": "a"}, goal="find flights", rules="one way only", operation="CLICK") + await _model(agent).decide_many(PICTURED, {"click_target": question}) + instructions = agent.calls[0][1]["click_target"]["instructions"].splitlines() + self.assertTrue(instructions[0].startswith("State: ")) + self.assertEqual(instructions[1:], ["Task: find flights", "one way only", "Which element should CLICK act on?"]) + + def test_a_long_page_text_does_not_push_out_the_recent_actions_or_the_elements(self) -> None: + state = { + "page": {"url": "https://flights", "title": "Flights", "text": "x" * 10_000}, + "elements": [{"index": "19", "role": "button", "label": "Search"}], + "recent_actions": [{"action": "Search", "kind": "click", "page_changed": False}], + } + context = omnijev_module.omnijev_context(Observation(state, images=(SCREEN,))) + self.assertLessEqual(len(context), omnijev_module.OMNIJEV_STATE_CHARS) + self.assertTrue(context.startswith('{"recent_actions":[{"action":"Search"')) + self.assertIn('"label":"Search"', context) + self.assertLess(context.count("x"), omnijev_module.OMNIJEV_PAGE_TEXT_CHARS + 10) + + def test_a_non_browser_state_is_sent_as_it_is(self) -> None: + context = omnijev_module.omnijev_context(Observation({"score": 1}, images=(SCREEN,))) + self.assertEqual(context, '{"score":1}') + + async def test_the_short_prompt_keeps_the_goal_and_the_ask_only(self) -> None: + agent = FakeMSO1() + question = ChoiceQuestion({"1": "a"}, goal="find flights", rules="one way only", operation="CLICK") + model = OmniJevModel(agent, model="m", prompt="short") + await model.decide_many(PICTURED, {"click_target": question}) + instructions = agent.calls[0][1]["click_target"]["instructions"] + self.assertEqual(instructions, "Task: find flights\nWhich element should CLICK act on?") + + def test_an_unknown_prompt_is_refused(self) -> None: + with self.assertRaises(ValueError): + OmniJevModel(FakeMSO1(), model="m", prompt="long") + + async def test_a_browser_element_row_becomes_its_label_and_value(self) -> None: + agent = FakeMSO1() + question = ChoiceQuestion( + {"12": {"element": "[12] Where from?", "current_value": "Zurich", "role": "combobox"}} + ) + await _model(agent).decide_many(PICTURED, {"click_target": question}) + self.assertEqual(agent.calls[0][1]["click_target"]["criteria"], {"12": "Where from? = Zurich"}) + + async def test_noul_criteria_are_spelled_out_after_the_question(self) -> None: + agent = FakeMSO1() + await _model(agent).decide_many(PICTURED, {"check": contract.CHECK}) + asked = agent.calls[0][1]["check"] + self.assertEqual(asked["type"], "noul") + self.assertTrue(asked["instructions"].endswith("true: the player stands\nfalse: the player hits")) + + async def test_a_choice_is_renormalised_over_the_offered_keys(self) -> None: + choice = (await _model().decide_many(PICTURED, {"pick": contract.PICK})).choice("pick") + self.assertEqual(choice.key, "stand") + self.assertAlmostEqual(sum(choice.probabilities.values()), 1.0, places=6) + self.assertAlmostEqual(choice.probabilities["stand"], 0.6 / 0.9, places=6) # abstain 0.1 left out + self.assertAlmostEqual(choice.confidence, (2 * 0.6 / 0.9 - 1) / 1, places=6) + + async def test_abstain_stays_in_the_raw_reply(self) -> None: + decision = await _model().decide_many(PICTURED, {"pick": contract.PICK}) + self.assertEqual(decision.raw["pick"]["abstain"], 0.1) + + async def test_a_noul_answer_is_the_probability_omnijev_gave(self) -> None: + noul = (await _model().decide_many(PICTURED, {"check": contract.CHECK})).noul("check") + self.assertEqual(noul.p, 0.3) + + +class TestFailures(IsolatedAsyncioTestCase): + async def test_an_observation_without_a_screenshot_is_a_config_error(self) -> None: + with self.assertRaises(BaseError) as caught: + await _model().decide_many(Observation({"page": "x"}), {"pick": contract.PICK}) + self.assertEqual(caught.exception.status, StatusCode.MODEL_SERVICE_CONFIG_ERROR) + self.assertIn("screenshot", str(caught.exception)) + + async def test_a_runtime_error_becomes_model_call_failed(self) -> None: + with self.assertRaises(BaseError) as caught: + await _model(FakeMSO1(error=RuntimeError("CUDA out of memory"))).decide_many( + PICTURED, {"pick": contract.PICK} + ) + self.assertEqual(caught.exception.status, StatusCode.MODEL_CALL_FAILED) + self.assertIn("CUDA out of memory", str(caught.exception)) + + +class TestFromEnv(TestCase): + def test_without_omnijev_repo_it_is_a_config_error_naming_the_variable(self) -> None: + with patch.dict(os.environ, {"OMNIJEV_REPO": ""}): + with self.assertRaises(BaseError) as caught: + OmniJevModel.from_env() + self.assertEqual(caught.exception.status, StatusCode.MODEL_SERVICE_CONFIG_ERROR) + self.assertIn("OMNIJEV_REPO", str(caught.exception)) + + def test_an_unknown_omnijev_prompt_is_a_config_error_before_any_load(self) -> None: + with patch.dict(os.environ, {"OMNIJEV_PROMPT": "long", "OMNIJEV_REPO": ""}): + with self.assertRaises(BaseError) as caught: + OmniJevModel.from_env() + self.assertEqual(caught.exception.status, StatusCode.MODEL_SERVICE_CONFIG_ERROR) + self.assertIn("OMNIJEV_PROMPT", str(caught.exception)) + + def test_a_folder_that_is_not_an_omnijev_clone_is_a_config_error(self) -> None: + with patch.dict(os.environ, {"OMNIJEV_REPO": str(Path(__file__).parent)}): + with self.assertRaises(BaseError) as caught: + OmniJevModel.from_env() + self.assertEqual(caught.exception.status, StatusCode.MODEL_SERVICE_CONFIG_ERROR) + + def test_the_default_checkpoint_is_the_cpu_sized_release(self) -> None: + self.assertEqual( + (omnijev_module.OMNIJEV_DEFAULT_CHECKPOINT, omnijev_module.OMNIJEV_DEFAULT_REVISION), + ("tinnel123/OmniJev-0.8B", "v1.1"), + ) diff --git a/uv.lock b/uv.lock index 52ec63f..7f51e25 100644 --- a/uv.lock +++ b/uv.lock @@ -13,6 +13,25 @@ resolution-markers = [ "python_full_version < '3.12' and sys_platform != 'emscripten' and sys_platform != 'win32'", ] +[[package]] +name = "accelerate" +version = "1.15.0" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "huggingface-hub" }, + { name = "numpy", version = "2.4.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.12'" }, + { name = "numpy", version = "2.5.3", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.12'" }, + { name = "packaging" }, + { name = "psutil" }, + { name = "pyyaml" }, + { name = "safetensors" }, + { name = "torch" }, +] +sdist = { url = "https://files.pythonhosted.org/packages/f5/b5/1d3ed029ac71d3f2961346829a268da923698e9fd63f218f78841f216bfd/accelerate-1.15.0.tar.gz", hash = "sha256:5654f8c5eaa0d4fa68b33e287a97765da6849bf6d51dcac874e73fbbddfb6134", size = 422615, upload-time = "2026-09-09T13:04:49.078Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/8a/4c/34f0450479d01195027260da68d8a3880683f1640c3ca5adf64acb3185f1/accelerate-1.15.0-py3-none-any.whl", hash = "sha256:97eacca0b73e45cb867dbf8c5d5d4dc32219544300e0c8992c7334dc2ef33cec", size = 394295, upload-time = "2026-09-09T13:04:47.331Z" }, +] + [[package]] name = "agentdescent" version = "0.4.6" @@ -3608,6 +3627,28 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/a2/9a/07d658e1e7fad860f1c541ab941348125dbdab773be3a0afaf32361866c7/pdfplumber-0.11.10-py3-none-any.whl", hash = "sha256:7741ea81bf165b474b153e6789d10d18e06b6ddcf3ec84289c3ef2fed6802580", size = 60047, upload-time = "2026-06-15T03:31:29.702Z" }, ] +[[package]] +name = "peft" +version = "0.21.0" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "accelerate" }, + { name = "huggingface-hub" }, + { name = "numpy", version = "2.4.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.12'" }, + { name = "numpy", version = "2.5.3", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.12'" }, + { name = "packaging" }, + { name = "psutil" }, + { name = "pyyaml" }, + { name = "safetensors" }, + { name = "torch" }, + { name = "tqdm" }, + { name = "transformers" }, +] +sdist = { url = "https://files.pythonhosted.org/packages/4f/91/56cc2b1b6824f5a4026750274240f26cd23465fb1a29cd14772d555433f2/peft-0.21.0.tar.gz", hash = "sha256:17f2b5a264439f4cd983c02e95954f2e948e3bcb85f75895ee307f94a55fd28f", size = 978334, upload-time = "2026-09-15T13:34:02.583Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/13/f0/29c37002f5ef5cb5be54890114d5c8f726f0cb5072a8ba0880758c912bfd/peft-0.21.0-py3-none-any.whl", hash = "sha256:b64eb75fd9dece7401c70e675b8d9de024993b70483691b41c62876f0c7809b7", size = 832883, upload-time = "2026-09-15T13:33:59.775Z" }, +] + [[package]] name = "pillow" version = "12.3.0" @@ -3958,6 +3999,34 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/e4/04/d52c7016b04b6c5108f26691f9d33ec82a9b65d041f1a9c771137693d618/protobuf-7.36.2-py3-none-any.whl", hash = "sha256:bdb3a345d48db958e6ce1f18e508beb0cc981d64f24088427549c866cd039f1e", size = 179806, upload-time = "2026-09-17T20:07:58.211Z" }, ] +[[package]] +name = "psutil" +version = "7.2.2" +source = { registry = "https://pypi.org/simple" } +sdist = { url = "https://files.pythonhosted.org/packages/aa/c6/d1ddf4abb55e93cebc4f2ed8b5d6dbad109ecb8d63748dd2b20ab5e57ebe/psutil-7.2.2.tar.gz", hash = "sha256:0746f5f8d406af344fd547f1c8daa5f5c33dbc293bb8d6a16d80b4bb88f59372", size = 493740, upload-time = "2026-01-28T18:14:54.428Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/51/08/510cbdb69c25a96f4ae523f733cdc963ae654904e8db864c07585ef99875/psutil-7.2.2-cp313-cp313t-macosx_10_13_x86_64.whl", hash = "sha256:2edccc433cbfa046b980b0df0171cd25bcaeb3a68fe9022db0979e7aa74a826b", size = 130595, upload-time = "2026-01-28T18:14:57.293Z" }, + { url = "https://files.pythonhosted.org/packages/d6/f5/97baea3fe7a5a9af7436301f85490905379b1c6f2dd51fe3ecf24b4c5fbf/psutil-7.2.2-cp313-cp313t-macosx_11_0_arm64.whl", hash = "sha256:e78c8603dcd9a04c7364f1a3e670cea95d51ee865e4efb3556a3a63adef958ea", size = 131082, upload-time = "2026-01-28T18:14:59.732Z" }, + { url = "https://files.pythonhosted.org/packages/37/d6/246513fbf9fa174af531f28412297dd05241d97a75911ac8febefa1a53c6/psutil-7.2.2-cp313-cp313t-manylinux2010_x86_64.manylinux_2_12_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:1a571f2330c966c62aeda00dd24620425d4b0cc86881c89861fbc04549e5dc63", size = 181476, upload-time = "2026-01-28T18:15:01.884Z" }, + { url = "https://files.pythonhosted.org/packages/b8/b5/9182c9af3836cca61696dabe4fd1304e17bc56cb62f17439e1154f225dd3/psutil-7.2.2-cp313-cp313t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:917e891983ca3c1887b4ef36447b1e0873e70c933afc831c6b6da078ba474312", size = 184062, upload-time = "2026-01-28T18:15:04.436Z" }, + { url = "https://files.pythonhosted.org/packages/16/ba/0756dca669f5a9300d0cbcbfae9a4c30e446dfc7440ffe43ded5724bfd93/psutil-7.2.2-cp313-cp313t-win_amd64.whl", hash = "sha256:ab486563df44c17f5173621c7b198955bd6b613fb87c71c161f827d3fb149a9b", size = 139893, upload-time = "2026-01-28T18:15:06.378Z" }, + { url = "https://files.pythonhosted.org/packages/1c/61/8fa0e26f33623b49949346de05ec1ddaad02ed8ba64af45f40a147dbfa97/psutil-7.2.2-cp313-cp313t-win_arm64.whl", hash = "sha256:ae0aefdd8796a7737eccea863f80f81e468a1e4cf14d926bd9b6f5f2d5f90ca9", size = 135589, upload-time = "2026-01-28T18:15:08.03Z" }, + { url = "https://files.pythonhosted.org/packages/81/69/ef179ab5ca24f32acc1dac0c247fd6a13b501fd5534dbae0e05a1c48b66d/psutil-7.2.2-cp314-cp314t-macosx_10_15_x86_64.whl", hash = "sha256:eed63d3b4d62449571547b60578c5b2c4bcccc5387148db46e0c2313dad0ee00", size = 130664, upload-time = "2026-01-28T18:15:09.469Z" }, + { url = "https://files.pythonhosted.org/packages/7b/64/665248b557a236d3fa9efc378d60d95ef56dd0a490c2cd37dafc7660d4a9/psutil-7.2.2-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:7b6d09433a10592ce39b13d7be5a54fbac1d1228ed29abc880fb23df7cb694c9", size = 131087, upload-time = "2026-01-28T18:15:11.724Z" }, + { url = "https://files.pythonhosted.org/packages/d5/2e/e6782744700d6759ebce3043dcfa661fb61e2fb752b91cdeae9af12c2178/psutil-7.2.2-cp314-cp314t-manylinux2010_x86_64.manylinux_2_12_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:1fa4ecf83bcdf6e6c8f4449aff98eefb5d0604bf88cb883d7da3d8d2d909546a", size = 182383, upload-time = "2026-01-28T18:15:13.445Z" }, + { url = "https://files.pythonhosted.org/packages/57/49/0a41cefd10cb7505cdc04dab3eacf24c0c2cb158a998b8c7b1d27ee2c1f5/psutil-7.2.2-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:e452c464a02e7dc7822a05d25db4cde564444a67e58539a00f929c51eddda0cf", size = 185210, upload-time = "2026-01-28T18:15:16.002Z" }, + { url = "https://files.pythonhosted.org/packages/dd/2c/ff9bfb544f283ba5f83ba725a3c5fec6d6b10b8f27ac1dc641c473dc390d/psutil-7.2.2-cp314-cp314t-win_amd64.whl", hash = "sha256:c7663d4e37f13e884d13994247449e9f8f574bc4655d509c3b95e9ec9e2b9dc1", size = 141228, upload-time = "2026-01-28T18:15:18.385Z" }, + { url = "https://files.pythonhosted.org/packages/f2/fc/f8d9c31db14fcec13748d373e668bc3bed94d9077dbc17fb0eebc073233c/psutil-7.2.2-cp314-cp314t-win_arm64.whl", hash = "sha256:11fe5a4f613759764e79c65cf11ebdf26e33d6dd34336f8a337aa2996d71c841", size = 136284, upload-time = "2026-01-28T18:15:19.912Z" }, + { url = "https://files.pythonhosted.org/packages/e7/36/5ee6e05c9bd427237b11b3937ad82bb8ad2752d72c6969314590dd0c2f6e/psutil-7.2.2-cp36-abi3-macosx_10_9_x86_64.whl", hash = "sha256:ed0cace939114f62738d808fdcecd4c869222507e266e574799e9c0faa17d486", size = 129090, upload-time = "2026-01-28T18:15:22.168Z" }, + { url = "https://files.pythonhosted.org/packages/80/c4/f5af4c1ca8c1eeb2e92ccca14ce8effdeec651d5ab6053c589b074eda6e1/psutil-7.2.2-cp36-abi3-macosx_11_0_arm64.whl", hash = "sha256:1a7b04c10f32cc88ab39cbf606e117fd74721c831c98a27dc04578deb0c16979", size = 129859, upload-time = "2026-01-28T18:15:23.795Z" }, + { url = "https://files.pythonhosted.org/packages/b5/70/5d8df3b09e25bce090399cf48e452d25c935ab72dad19406c77f4e828045/psutil-7.2.2-cp36-abi3-manylinux2010_x86_64.manylinux_2_12_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:076a2d2f923fd4821644f5ba89f059523da90dc9014e85f8e45a5774ca5bc6f9", size = 155560, upload-time = "2026-01-28T18:15:25.976Z" }, + { url = "https://files.pythonhosted.org/packages/63/65/37648c0c158dc222aba51c089eb3bdfa238e621674dc42d48706e639204f/psutil-7.2.2-cp36-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:b0726cecd84f9474419d67252add4ac0cd9811b04d61123054b9fb6f57df6e9e", size = 156997, upload-time = "2026-01-28T18:15:27.794Z" }, + { url = "https://files.pythonhosted.org/packages/8e/13/125093eadae863ce03c6ffdbae9929430d116a246ef69866dad94da3bfbc/psutil-7.2.2-cp36-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:fd04ef36b4a6d599bbdb225dd1d3f51e00105f6d48a28f006da7f9822f2606d8", size = 148972, upload-time = "2026-01-28T18:15:29.342Z" }, + { url = "https://files.pythonhosted.org/packages/04/78/0acd37ca84ce3ddffaa92ef0f571e073faa6d8ff1f0559ab1272188ea2be/psutil-7.2.2-cp36-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:b58fabe35e80b264a4e3bb23e6b96f9e45a3df7fb7eed419ac0e5947c61e47cc", size = 148266, upload-time = "2026-01-28T18:15:31.597Z" }, + { url = "https://files.pythonhosted.org/packages/b4/90/e2159492b5426be0c1fef7acba807a03511f97c5f86b3caeda6ad92351a7/psutil-7.2.2-cp37-abi3-win_amd64.whl", hash = "sha256:eb7e81434c8d223ec4a219b5fc1c47d0417b12be7ea866e24fb5ad6e84b3d988", size = 137737, upload-time = "2026-01-28T18:15:33.849Z" }, + { url = "https://files.pythonhosted.org/packages/8c/c7/7bb2e321574b10df20cbde462a94e2b71d05f9bbda251ef27d104668306a/psutil-7.2.2-cp37-abi3-win_arm64.whl", hash = "sha256:8c233660f575a5a89e6d4cb65d9f938126312bca76d8fe087b947b3a1aaac9ee", size = 134617, upload-time = "2026-01-28T18:15:36.514Z" }, +] + [[package]] name = "py-key-value-aio" version = "0.4.6" @@ -5188,6 +5257,16 @@ dev = [ laya = [ { name = "laya" }, ] +omnijev = [ + { name = "accelerate" }, + { name = "huggingface-hub" }, + { name = "peft" }, + { name = "pillow" }, + { name = "safetensors" }, + { name = "torch" }, + { name = "torchvision" }, + { name = "transformers" }, +] report = [ { name = "pillow" }, { name = "playwright" }, @@ -5195,15 +5274,19 @@ report = [ [package.metadata] requires-dist = [ + { name = "accelerate", marker = "extra == 'omnijev'" }, { name = "ai2thor", marker = "extra == 'alfworld-visual'", specifier = "==2.1.0" }, { name = "alfworld", marker = "extra == 'alfworld'", specifier = ">=0.4" }, { name = "cua-s1", marker = "extra == 'cua'", git = "https://github.com/trycua/cua.git?subdirectory=libs%2Fcua-s1%2Fpython&rev=aea61b6eb97e2d8c0f6f71eb804e5769fe910af4" }, { name = "httpx", specifier = ">=0.28" }, { name = "huggingface-hub", marker = "extra == 'cua'", specifier = ">=0.24" }, + { name = "huggingface-hub", marker = "extra == 'omnijev'", specifier = ">=0.24" }, { name = "laya", marker = "extra == 'laya'", specifier = ">=0.3.4" }, { name = "mcp", specifier = ">=1.26" }, { name = "opencv-python-headless", marker = "extra == 'alfworld-visual'", specifier = ">=4.10" }, { name = "openjiuwen", git = "https://github.com/ThinkFlowLab/agent-core?rev=jj-0.1.0" }, + { name = "peft", marker = "extra == 'omnijev'", specifier = ">=0.15" }, + { name = "pillow", marker = "extra == 'omnijev'" }, { name = "pillow", marker = "extra == 'report'", specifier = ">=10" }, { name = "playwright", marker = "extra == 'report'", specifier = ">=1.45" }, { name = "pytest", marker = "extra == 'dev'", specifier = ">=8" }, @@ -5211,14 +5294,18 @@ requires-dist = [ { name = "python-dotenv", specifier = ">=1.0" }, { name = "rlcard", marker = "extra == 'blackjack'", specifier = ">=1.2" }, { name = "ruff", marker = "extra == 'dev'", specifier = ">=0.6" }, + { name = "safetensors", marker = "extra == 'omnijev'" }, { name = "system1-agents", extras = ["alfworld"], marker = "extra == 'alfworld-visual'" }, { name = "textworld", marker = "extra == 'alfworld'", specifier = ">=1.7" }, { name = "torch", marker = "extra == 'alfworld-visual'", specifier = ">=2" }, + { name = "torch", marker = "extra == 'omnijev'", specifier = ">=2.4" }, { name = "torchvision", marker = "extra == 'alfworld-visual'", specifier = ">=0.15" }, + { name = "torchvision", marker = "extra == 'omnijev'", specifier = ">=0.15" }, + { name = "transformers", marker = "extra == 'omnijev'", specifier = ">=5.0" }, { name = "ty", marker = "extra == 'dev'", specifier = ">=0.0.83" }, { name = "werkzeug", marker = "extra == 'alfworld-visual'", specifier = "==2.0.3" }, ] -provides-extras = ["blackjack", "alfworld", "alfworld-visual", "report", "laya", "cua", "dev"] +provides-extras = ["blackjack", "alfworld", "alfworld-visual", "report", "laya", "cua", "omnijev", "dev"] [[package]] name = "tatsu"