feat(skills): atlan-search skill for the Claude Code and Cursor plugins - #236
ankitjaggi wants to merge 4 commits into
Conversation
…ugins Transposes the discovery judgment from Wisdom's `category: discover` skills onto the MCP tool surface. Wisdom's versions are SQL over the `gold.*` lakehouse tables; the SQL does not carry over, but the judgment does — which tool answers which ask, and the filter rules that decide whether it returns the right thing. One skill, six references (progressive disclosure), replacing twelve SQL-shaped ones. What survived the transposition: - tags are the `tags` parameter, not a `conditions` entry - custom metadata needs `resolve_metadata` before it can be filtered - owner fields hold usernames, not display names — `get_users` first - a table's columns come from `get_assets(attributes=["columns"])` - lineage needs the right instance picked from all same-named hits - a domain's assets need the subdomain/product walk, not just the root GUID - DQ pass/fail values differ per source AND differ in case - glossary qualifiedNames are opaque, so scope via glossary_qualified_name - asset-type precision, and the internal-type exclusions when going broad `cursor-plugin/skills` is a symlink to the repo-root `skills/`, so a skill is authored once and ships to both plugins. Not yet measured: the A/B run that justifies this lands with the eval dataset in agent-toolkit-internal. Rules that duplicate what the tool descriptions already say should be cut from the skill rather than shipped — the diff is what tells us which those are. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
@ankitjaggi thanks for putting this together. I checked the references against the latest agent-toolkit-internal, then tested every finding against a live tenant. Adding a couple of comments below (Also – your open question 2 looks fine. namespace_type="classification" works. middleware/arg_normalization.py:_resolve_enum_value maps singular to plural before validation runs. Nothing to change.) 1.
|
…ated rules Review findings from #236: - lineage.md: traverse_lineage defaults to immediate_neighbors=true, which makes depth a no-op and silently returns a one-hop graph with has_more=false. Document the default, say when to set immediate_neighbors=false, and use the canonical param name limit (default 10, max 20) instead of the legacy size alias. - filters.md: directly_tagged defaults to true (direct tags only), the opposite of what the reference claimed. Flipped, with the guidance to pass false for governance questions that need propagated tags. - SKILL.md: drop rules 1, 2 and 11 — near-verbatim restatements of the semantic_search_tool user_query description and docstring. Kept the count-vs-list rule, whose substance is not in the tool schema. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
@benhuds thanks — both findings reproduce against 1. One correction on the param name: 2. Open question 1 — cut 1, 2, 11; kept 10. Agreed on 1, 2 and 11: they restate the Worth flagging: the A/B eval that was meant to measure these rules never ran (blocked on eval venv + proxy access), so cutting those three is a judgment call on duplication, not a measured delta. If you'd rather have the numbers before merge, I can hold on that part. Thanks for checking question 2 on |
Analysed 30 days of the `mcp` Braintrust project, scoped to the clients this
skill ships to (claude-code + cursor-vscode): search_assets_tool 11,260 calls,
traverse_lineage_tool 5,406, semantic_search_tool 2,479.
SKILL.md
- Call-shape rule (new 10). metadata.correction_guidance over 30d shows the
top rejects: `query` for `user_query` (53), `conditions` as a string (26),
`filters` (24), `limit` out of range (24), `direction: "BOTH"` (10),
`type_names`/`asset_types` (10). None of this was covered.
- Compound-ask rule (new 11) and non-catalog-ask rule (new 12): median query
is 7 words, p90 13, and asks routinely bundle resolve + owners + tags.
- Columns rule now covers the reverse direction (which table has column X) —
a recurring ask the forward-only rule left unanswered.
- Naming rule now lists the attributes to request: displayName appeared in 2
of 276 calls against name's 61, and 53% of calls request no attributes at
all, which made the rule unactionable as written.
- Popularity rule distinguishes "biggest" (rowCount) from "most used".
- Routing table: counts go to search_assets (return_count_only /
aggregations), not semantic_search; added a data-quality row.
filters.md
- Full operator table. `within`, `endswith`, `between`, `regexp`, `fuzzy`,
`has_any_value` were all undocumented.
- The `{"name": ..., "name2": ...}` anti-pattern: 42 calls across 11 distinct
clients in 30 days, every one a hard error. A JSON object cannot repeat a
key; the answer is a `within` list.
- "Missing a description" via negative_conditions + has_any_value.
- Canonical single-table resolution (name + databaseName + schemaName) and
splitting a pasted db.schema.table instead of sending it as prose.
data-quality.md
- dqRuleTemplateName plus the observed identifiers. Traces show invented
compounds (CUSTOM_SQL_QUERY, CUSTOM_SQL_ROW_COUNT) and UI labels
("Standard Deviation") passed as filter values.
lineage.md
- direction has no BOTH; 10 rejected calls in 30 days.
Sample values are synthetic; no tenant or customer identifiers included.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
What 30 days of production traces say the skill is missingPulled the First, a correction to how I'd been framing the tool mix. The
1. Call shapes — biggest measurable failure class, previously uncovered
~150 wasted calls. New rule 10 covers all of it; 2.
|
A/B eval results — the numbers benhuds asked forRan the golden set twice against The new trace-derived cases: 4/14 → 9/14, no regressionsFixed: Across three independent baseline runs the baseline scored 6, 6, 4 of 14; the skills arm scored 10, 10, 9. Individual case attribution is noisy at n=14, but the aggregate separation is consistent. The original cases: 7/14 → 8/14This gap is the most useful thing in the run. The original dataset was written to guard the rules the skill had before review — including the four you identified as near-verbatim from the Its one "regression" is a case-design bug, not a skill regression. Honest limitsLLMJudge is flat, in both datasets. The skill decisively improves how the agent calls the tools — right tool, right arguments, right order — and shows no measurable improvement in final answer quality as scored by a blind judge. I am not going to dress that up: on this tenant, for these 28 asks, the mechanical gains did not translate into answers the judge rated higher. Worth knowing before anyone claims the skill makes answers better. I burned two runs on my own bad rubrics. My first pass wrote mechanism into Rule 12 still fails. Pre-existing cases worth a separate look (not touched, to avoid retuning rubrics mid-measurement): Infrastructure noteConcurrency 16 saturates the inference proxy — connect errors, agent and judge timeouts, and judge scores coming back as "failed to parse". Every run at concurrency 4 completed with zero infrastructure errors. The A/B in the Makefile should probably drop its default concurrency for the skills arm, since a 27k-char system prompt on every case is what tips it over. |
|
@Aryamanz29 to complement these lets also add a line in server instructions on using these skills in case the user has them? |
New rule 7, from a customer evaluation where the same scenario passed with a real business-unit value and failed on all three models with a plausible but non-existent one. Rule 6 already says tags and custom metadata are not `conditions` — the slot. Nothing covered the value. A wrong value fails silently: the search returns an empty set, which the agent reports as "nothing is tagged that way". In the customer's traces the models cycled through eight invented business-unit values across three models and never once called resolve_metadata to list the real ones. The same scenario with a correct value passed 3/3. The rule: resolve_metadata first, filter on a value from that list, and if nothing matches report which values do exist rather than answering zero. Plus the reason it matters — "no PII assets in that business unit" is a dangerous sentence when the truth is that the unit name was misspelled. Example values are synthetic; renumbers former rules 7-12 to 8-13. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Customer evaluation traces — one new rule, and two product bugs a skill can't fixA customer ran 19 scenarios against the MCP across Haiku/Sonnet/Opus (11/19, 14/19, 15/19). I pulled their tenant's traces for the last 30 days — 1,157 spans — and read the actual tool calls behind each failure. Tenant and asset names stay out of this repo; the example values in the diff are synthetic. New rule 7: resolve tag and custom-metadata values before filtering (
|
What
Adds
skills/atlan-search/— the first agent skill shipped through the Atlan plugins — plus the wiring that gets it to both Claude Code and Cursor.The content is a transposition of Wisdom's
category: discoverskills (gold-layer-assets,-glossary,-tags,-custom-metadata,-owners,-lineage,-data-mesh,-data-quality,-bi-assets,-pipelines). Those are written as SQL over thegold.*lakehouse tables. The SQL does not transpose tosemantic_search/search_assets— the judgment does, and that is what is here.Layout
cursor-plugin/skillsis a symlink to the repo-rootskills/so a skill is authored once and ships to both plugins. Claude Code discoversskills/at the plugin root (source: "./"); no manifest change needed.What carried over from Wisdom
array_contains(tags, …)doesn't exist → jointagrelationshiptagsparameter, not aconditionsentrycustommetadatarelationshipresolve_metadatafirst, then the set's display name inattributesowner_usersholdsjohn.doe, not "John Doe"get_usersfirst, then filter;any_conditionsfor user-or-grouptable_columns[]GUID expansionget_assets(attributes=["columns"])has_lineage = truetraverse_lineagedomain_guidsPASSvs lowercase partner trap{nanoId}@{glossaryQN}glossary_qualified_name, never pattern-match the QNasset_typefilter + 63 noise typesValidation
This has not been A/B'd yet. The eval dataset and the harness that measures it are in
agent-toolkit-internal(companion PR): asearch_skillscapability with one case per rule this skill teaches, andcli.py run --with-skills/cli.py compareto diff the arms case-by-case.Run before merge:
Open questions for review
semantic_search/search_assetsdocstrings. If the A/B shows no movement on those cases, the right fix is to cut them from the skill, not ship them — the skill should carry what the tool surface cannot.resolve_metadatanamespace values.references/filters.mdusesnamespace_type="classification"and"business_metadata". The latter is confirmed by the tool's own docstring; the former is inferred and unverified against the parameter's accepted values. Please confirm or correct.gold-layer-assets.md, where it exists becausegold.assetsis a partitioned SQL table. Whether persona-scopedsearch_assetsresults surface those types at all is untested — if they never appear, the list is dead weight in the context window and should be trimmed.data-qualityskill on its own), that is a different split.Unrelated staleness noticed and deliberately left alone:
CLAUDE.mddescribes 15 tools (the surface is 41), andREADME.mdlinks a non-existentclaude-plugin/directory.🤖 Generated with Claude Code