Summary
Same defect as packages/llm/llm-anthropic/src/anthropic-adapter.ts and packages/llm/llm-openai/src/openai-adapter.ts (filed separately), present in the Gemini adapter:
const contextLimit = getContextLimit(this.provider, this.model);
const { evidence: trimmedEvidence, trim } = trimEvidenceForContext(evidence, contextLimit);
const response = await withRetry(() =>
this.client.models.generateContent({
model: this.model,
contents: buildUserPrompt(trimmedEvidence, question), // model sees trimmedEvidence
...
}),
);
...
const grounding = validateEvidenceGrounding(evidence, evidenceRefs);
// ^^^^^^^^ should be trimmedEvidence
Why this matters
The prompt sent to Gemini (contents: buildUserPrompt(trimmedEvidence, question)) is built from the trimmed evidence object, which trimEvidenceForContext (packages/llm/llm-core/src/context.ts) may have shrunk by dropping evidence items, what_changed entries, or similar-incident entries to stay under the model's context budget. But the hallucination check afterward, validateEvidenceGrounding(evidence, evidenceRefs), is run against the original, untrimmed evidence.
Because validateEvidenceGrounding (packages/llm/llm-core/src/grounding.ts) only checks id membership against evidence.evidence, and trimmedEvidence.evidence is a subset of that, an evidence_refs entry that references an evidence item the model was never actually shown (because it was trimmed) can still be marked as "grounded". This defeats the purpose of grounded_evidence_refs / ungrounded_evidence_refs / groundedness_warning, exactly in the large-trace case where trimming (and therefore the risk of the model's answer drifting from what it was shown) is most likely.
Suggested fix
const grounding = validateEvidenceGrounding(trimmedEvidence, evidenceRefs);
Files:
packages/llm/llm-gemini/src/gemini-adapter.ts
Summary
Same defect as
packages/llm/llm-anthropic/src/anthropic-adapter.tsandpackages/llm/llm-openai/src/openai-adapter.ts(filed separately), present in the Gemini adapter:Why this matters
The prompt sent to Gemini (
contents: buildUserPrompt(trimmedEvidence, question)) is built from the trimmed evidence object, whichtrimEvidenceForContext(packages/llm/llm-core/src/context.ts) may have shrunk by dropping evidence items,what_changedentries, or similar-incident entries to stay under the model's context budget. But the hallucination check afterward,validateEvidenceGrounding(evidence, evidenceRefs), is run against the original, untrimmedevidence.Because
validateEvidenceGrounding(packages/llm/llm-core/src/grounding.ts) only checks id membership againstevidence.evidence, andtrimmedEvidence.evidenceis a subset of that, anevidence_refsentry that references an evidence item the model was never actually shown (because it was trimmed) can still be marked as "grounded". This defeats the purpose ofgrounded_evidence_refs/ungrounded_evidence_refs/groundedness_warning, exactly in the large-trace case where trimming (and therefore the risk of the model's answer drifting from what it was shown) is most likely.Suggested fix
Files:
packages/llm/llm-gemini/src/gemini-adapter.ts