Skip to content

grammar: clone finds each stack element's rule by binary search - #336

Open
professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:grammar-clone-bsearch
Open

professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:grammar-clone-bsearch

Conversation

@professorpalmer

Copy link
Copy Markdown

Overview

llama_grammar_clone_impl points each stack element of the copy at the copied rules. It found each element with a search over every element of every rule: O(stack elements x rule elements) for each clone. The server clones the sampler, and with it the grammar, on each speculative step (common_sampler_clone before common_sampler_sample_and_accept_n in server-context.cpp). With a tool list, the tool-call grammar is large, so this cost is in every decode step.

This change finds each element's rule with a binary search on the rule start addresses (std::less on the pointers, then the offset inside the rule). The copy itself does not change.

Additional information

  • Clone cost, the server's tool-call grammar for 107 tools (Bonsai chat template, 266k-character grammar: 4,704 rules, 43,790 elements, 109 stacks, 323 stack elements): one clone 8.4 ms -> 0.40 ms (CPU, mean of 50, two passes). The old and the new clone map every stack element to the same rule and offset.
  • Serving, RTX 4070, Bonsai 2 27B PTQ1_0 with the MTP head, 107 tools (106.7k characters of tool definitions), about 63k tokens, our fork with the same change: plan-like answer 41.8 -> 61.2 tok/s; the same greedy 47.6 -> 69.9; a request with three tool calls 60.1 -> 87.7. Each pair has the same token count and draft acceptance, and the tool-call request gives the same calls (same hash). For comparison, 2 tools at the same depth: 63.5 tok/s. Found from a user report: Real tests on pi-agent, slow (in the last build and in all xD) professorpalmer/bonsai-ada-surgery#7 (Pi agent with 107 MCP tools).
  • Test: test-grammar-integration gets a clone check. It clones at each position of a JSON-schema input, frees the original, checks that every stack element of the clone points into the clone's own rules, and checks that the clone matches the rest of the input. Negative control: with the redirect removed, the test fails at the pointer check. test-grammar-parser, test-llama-grammar and test-grammar-integration pass (static CPU build, Windows, MSVC).
  • Upstream ggml-org has the same loop in src/llama-grammar.cpp and the same per-step clone in the server.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. An AI coding assistant (Claude), working for the account owner at the owner's request, traced the slow decode step from the serving A/B to this loop, wrote the change and the test, and ran the timing, the unit tests, the negative control and the serving A/B.

llama_grammar_clone_impl redirected each stack element to the copied rules with a search over every element of every rule, O(stack elements x rule elements) for each clone. The server clones the sampler, and with it the grammar, on each speculative step, so with a large tool-call grammar this cost is in every decode step. Now each element's rule is found by a binary search on the rule start addresses; the copy itself does not change.

Server tool-call grammar for 107 tools (4704 rules, 43790 elements, 323 stack elements): one clone 8.4 ms -> 0.40 ms (CPU). test-grammar-integration: clone at each position of an input, check that the clone's stacks point into its own rules and that the clone matches the rest of the input.

Reported-by: Milor123 (professorpalmer/bonsai-ada-surgery#7)
@professorpalmer

Copy link
Copy Markdown
Author

Confirmed by the user who reported it (professorpalmer/bonsai-ada-surgery#7), RTX 4070, Pi agent with 107 MCP tools, same chat at 64-71k: 45.9 -> 63.3 tok/s with this change; time per draft step 44.8 -> 35.7 ms.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant