Skip to content

Refactor CLI tools around one search - #76

Open
yorkeccak wants to merge 3 commits into
mainfrom
refactor/simplify-cli-tools
Open

yorkeccak wants to merge 3 commits into
mainfrom
refactor/simplify-cli-tools

Conversation

@yorkeccak

@yorkeccak yorkeccak commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Refactors the CLI's tools into a smaller surface, in which each command has one clear job:

Command Job
valyu search <query> One search across the live web and every specialised dataset, routed automatically
valyu sources [need] Find the dataset id to scope a search with (ranked by what you need, or the full catalog)
valyu contents <url>... Read URLs as clean markdown or a summary
valyu deepresearch ... Start, check, steer, answer, share, cancel or delete an async research task

Nothing about scripting changes: JSON when piped or with -q, stdin input, {"error":{"message","code"}} on stderr, exit code 1 on error.

Why

The default search never reached the specialised datasets. With no type given, valyu search "<q>" sent search_type: "web". The same query ("GLP-1 receptor agonists cardiovascular outcomes" -n 5):

  • main: results_by_source: {"web": 5, "proprietary": 0}, first result 24,991 characters
  • this branch: web, PubMed and medRxiv results mixed by relevance, each capped at 4,000 characters, at the same price per result

The search types overlapped and went stale. paper / bio / finance / sec / patent / economics / news each hardcoded a list of sources. A question about a biotech's trial pipeline legitimately matches four of them, and a fixed list can't track a catalog that keeps changing. Routing already picks corpora per query, so the types are no longer part of the surface (old scripts still work, see Compatibility) and scoping is an explicit, validated option.

Output was larger than it needed to be:

Output Before After
Search content per result (default) up to 25,000 chars (API default) 4,000 chars (-l to change)
valyu sources -q ~165 KB (63% of it per-dataset response schemas) 69 KB
deepresearch status -q on a sample completed task ~46 KB ~30 KB (the agent's messages transcript is dropped)

What changed

search

  • One query argument; unquoted words are joined, and stdin works as before.
  • No search_type is sent, so an unscoped search covers the web plus every dataset on the plan and can never fail on access.
  • Advanced options: --include-source, --exclude-source, --source-bias, --start-date, --end-date, --country.
  • Scope values are checked before searching, because one unknown value fails the whole request upstream. They're checked against the live catalog plus the datasets the plan lists, since some searchable ids (e.g. valyu/valyu-fedwatch) aren't in the catalog.
    • A bare name like pubmed is rejected with did you mean: valyu/valyu-pubmed? rather than silently rewritten.
    • Presets in --exclude-source are refused, because the API doesn't expand them there.
  • -n accepts 1-100. -l, --response-length takes a character count (500-100000, default 4000). Text content is held to it locally, since some corpora return longer chunks than requested. Structured records (e.g. clinical trials, which arrive as JSON strings) are left whole so they still parse.
  • JSON gains a top-level hint only when the results need explaining:
    • a backend failed (retry, don't reword)
    • nothing matched
    • results carry no data (almost always a date filter on a structured dataset)
    • results were trimmed
    • a scoped dataset is outside the plan
  • The terminal view shows source, date and relevance for each result.

sources

  • valyu sources "<the data you need>" ranks the catalog and shows each dataset's example queries. It falls back to local ranking if the ranking endpoint is unavailable.
  • valyu sources lists the catalog grouped by category. Datasets outside the plan are marked LOCKED (locked in JSON), and the plan name is shown.
  • JSON drops display-only fields and per-dataset response schemas.

contents

  • --summary [instructions], -l, --response-length <chars> (default 30000), --extract-effort, --screenshot. At most 10 URLs; URLs can be piped in.
  • --summary https://... no longer swallows the URL as the instruction (this was already broken on main).
  • The terminal view prints the extracted text (it used to show only the character count), and cost is read from total_cost_dollars.
  • hint appears when nothing could be extracted or pages were truncated.

deepresearch

  • create:
    • The brief can be piped in.
    • --mode defaults to fast.
    • A PDF is opt-in with --pdf.
    • Adds --charts.
    • Adds --workflow <slug> -P key=value [--workflow-version n] to run a saved template (the template's own mode applies unless --mode is passed). Values with leading zeros (a CIK, 0700) stay strings.
    • Help text says when to use deep research rather than a few searches.
  • status [id]:
    • With an id, it's instant: a status line while running, the full report once complete. A paused task's JSON includes a hint with the exact respond payload.
    • Without an id, it lists recent tasks.
    • --wait <s> blocks up to that long and returns early when the task finishes or pauses at a checkpoint.
  • watch:
    • Piped or with -q, a task paused at a checkpoint returns its status with a hint naming the exact respond payload. It used to fall into interactive prompts.
    • It survives up to five consecutive transient failures (5xx, 429, network). A single gateway timeout used to end a multi-hour watch.
    • The timeout now covers max mode.
  • steer <id> <instruction> replaces update (kept as an alias).
  • share <id> publishes and share <id> --off unpublishes. It used to toggle, which isn't safe to repeat.
  • delete outside a terminal requires --yes instead of waiting on a prompt.
  • cancel, steer, delete and share now print JSON in JSON mode.
  • Fixed the interactive checkpoint replies:
    • planning questions now send answers as [{question, answer}], the documented shape (it was an object keyed q0, q1, ...)
    • source review now sends included_domains alongside excluded_domains

Errors

API errors keep their http_<status> codes, and the message now says what to do next:

  • 402: out of credits, so a retry won't help
  • 403 for a dataset above the plan: retry unscoped
  • 403 for sources search can't reach: named in the message
  • 429: wait and retry
  • 5xx: usually transient
  • 401, or a bare 403 Forbidden: re-login

A 403 that is a limit (e.g. more than 20 results) is passed through as-is instead of being blamed on the key.

Kept, but out of the main help

answer, batch and workflows still work unchanged but are hidden from valyu --help and the agent skill. Answers are best written from search results, templates now run through deepresearch create --workflow, and batches are many deepresearch tasks. Whether to remove them or bring them back is a follow-up decision.

Docs

  • SKILL.md is rewritten around the four commands. It covers when to scope, query-writing rules, reading hint, and when not to start deep research.
  • references/search.md, contents.md, deepresearch.md and error-codes.md are updated, and references/sources.md is new.
  • answer.md and workflows.md are removed from the skill.
  • The README is updated.

Compatibility

Scripts written against earlier versions keep working after an upgrade: every old form is still accepted, just no longer shown in --help or the skill. The compatibility code lives in src/lib/compat.ts.

Still accepted, and sending the same request as before:

  • valyu search <type> <query>: all eight types. Same fixed scope and default length as before, plus a note on stderr.
  • Old search flags: --max-price, --relevance-threshold, --search-type, --instructions, --fast-mode, --url-only, --no-tool-call.
  • Named lengths: short / medium / large / max for -l on search and contents.
  • Old contents flags: --length, --max-price-dollars, --structured, --structured-file.
  • Old deepresearch create flags: --output-format, --no-pdf, --alert-email-url.
  • Other commands: deepresearch update (alias of steer), sources list, sources categories, and answer / batch / workflows.

I checked this by running v1.2.3 and this branch side by side on 27 old-style invocations with fetch mocked, and diffing the request bodies. Every one exits 0, and all of them send the same request except these deliberate default changes:

Invocation with no options Before Now
valyu search "<q>" search_type: web, API default length unscoped (web + datasets, same price per result in spot checks), 4,000 chars per result
valyu contents <url> response_length: medium (50,000) 30,000
valyu deepresearch create "<brief>" standard, markdown + PDF fast, markdown

Behaviour that did change:

  • share publishes rather than toggles. Use share --off to unpublish.
  • Delete needs --yes outside a terminal. deepresearch delete there requires --yes instead of waiting on a prompt.
  • Status JSON drops the transcript. deepresearch status / watch no longer include messages.
  • sources categories returns the catalog listing. It used to return the bare category list.
  • Async contents is removed. contents --async / --watch / --webhook-url and contents jobs are gone, and a call takes at most 10 URLs.

Decisions made along the way

  • Hide, don't delete. Old commands and flags are hidden rather than removed, so no existing script breaks on upgrade.
  • Old search types keep their exact old scope. The type in valyu search paper "..." is a scope the caller chose, so it's honoured exactly. The new surface just stops advertising it.
  • hint is additive. It's absent when there's nothing to say, so jq '.results' pipelines are unaffected.
  • Content is clipped in JSON too. --response-length is honoured in the JSON output (except for structured records), so the flag means what it says.
  • Deep research --include-source is not validated locally. Deep research can use datasets that plain search can't reach, so a search-side check would reject valid ids.

Test plan

  • pnpm typecheck, pnpm test (82 tests; new: source resolution, ranking, plan coverage, hints, clipping, parsers, error messages, the non-interactive checkpoint path, transcript stripping), pnpm build
  • Live against the API:
    • search: unscoped, dataset-scoped (SEC filings with metadata, clinical trials, valyu-fedwatch), domain-scoped, stdin, old typed forms, old flags and named lengths
    • scope validation: bare name, typo and preset-in-exclude all rejected with suggestions
    • option validation errors, and a bogus key
    • terminal rendering
  • Live sources: ranked (semantic), full listing, --category, sources list, sources categories, terminal rendering
  • Live contents: full text, --summary with an instruction, --summary <url>, stdin, a 404 URL, the 11-URL limit, and the old --length / named sizes / --max-price-dollars
  • Old-vs-new request diff with fetch mocked: 27 old-style invocations across all commands, including hidden answer / workflows run / batch create
  • Live deepresearch:
    • create (piped brief, --metadata)
    • status with and without an id
    • steer, update alias, cancel, watch to completion
    • terminal report rendering
    • delete guard without --yes
    • validation errors
  • Not exercised live:
    • share, because it publishes a public link
    • delete --yes, because it's irreversible
    • LOCKED marks, because the test key is on the top plan (unit-tested)
    • the paused-checkpoint path through watch, because a fast task with --hitl plan-review completed without pausing (unit-tested)

Collapse the per-vertical search types into a single `search` that sends no
search_type, so an unscoped query routes across the web and every dataset
the plan covers. Scoping becomes explicit and validated: --include-source and
--exclude-source are checked against the live catalog (unknown ids come back
with suggestions), alongside --source-bias, date, country, -n 1-100 and a
character-count --response-length (default 4000, held locally).

Add `sources [need]` to rank the catalog for the data needed, with example
queries and plan coverage. Slim `contents` to summary, response length,
extract effort and screenshot. Align `deepresearch`: fast by default, PDF
opt-in, --workflow templates, `status [id] --wait`, `steer`, explicit
`share --off`, a --yes guard on delete outside a terminal, checkpoint
hand-back when `watch` is not interactive, retries on transient status
errors, and no agent transcript in status output. Fix the interactive
checkpoint reply shapes for planning questions and source review.

API error messages now say what to do next while keeping http_<status>
codes. Search and contents JSON carry a `hint` only when results need
explaining. `answer`, `batch` and `workflows` stay callable but leave the
main help. The agent skill, references and README are rewritten to match.
Treat a leading search type as legacy only when the form is unambiguous (a
quoted or piped query), so `valyu search news today` keeps both words.
Accept dataset ids the plan lists even when the catalog does not, and
suggest them for typos. Leave structured records whole when holding content
to --response-length, since trials arrive as JSON strings. Give a bare 403
Forbidden the rejected-key advice, keep key=value values with leading zeros
as strings, stop --summary from swallowing a following URL, add the
checkpoint hint to `deepresearch status` JSON, report unreadable --file
paths plainly, and fix a help-text line continuation.
Scripts written against earlier versions should not break on upgrade, so
the old forms stay accepted without appearing in --help or the skill:

- `valyu search <type> <query>` sends exactly the request it used to, with
  the same fixed scope and default length, and prints a note to stderr.
- --max-price, --relevance-threshold, --search-type, --instructions,
  --fast-mode, --url-only and --no-tool-call pass through as before.
- Named lengths (short, medium, large, max) work for search and contents,
  and contents keeps --length, --max-price-dollars and --structured.
- deepresearch create keeps --output-format, --no-pdf and --alert-email-url.
- `sources categories` lists the catalog.

The compatibility code lives in src/lib/compat.ts.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant