Skip to content

feat(skills): atlan-search skill for the Claude Code and Cursor plugins - #236

Open
ankitjaggi wants to merge 4 commits into
mainfrom
feat/atlan-search-skill
Open

ankitjaggi wants to merge 4 commits into
mainfrom
feat/atlan-search-skill

Conversation

@ankitjaggi

Copy link
Copy Markdown
Collaborator

What

Adds skills/atlan-search/ — the first agent skill shipped through the Atlan plugins — plus the wiring that gets it to both Claude Code and Cursor.

The content is a transposition of Wisdom's category: discover skills (gold-layer-assets, -glossary, -tags, -custom-metadata, -owners, -lineage, -data-mesh, -data-quality, -bi-assets, -pipelines). Those are written as SQL over the gold.* lakehouse tables. The SQL does not transpose to semantic_search / search_assets — the judgment does, and that is what is here.

Layout

skills/atlan-search/
├── SKILL.md                        tool routing table + 12 universal rules
└── references/
    ├── filters.md                  search_assets cookbook, type lists, noise exclusions
    ├── lineage.md                  instance resolution, then traversal
    ├── glossary.md                 terms, categories, hierarchy, linked assets
    ├── data-products.md            domains, subdomains, products, ports, lifecycle
    ├── data-quality.md             native DQ + Soda / Anomalo / Monte Carlo
    └── bi-and-pipelines.md         per-vendor BI types; Airflow, dbt, Fivetran, ADF

cursor-plugin/skills is a symlink to the repo-root skills/ so a skill is authored once and ships to both plugins. Claude Code discovers skills/ at the plugin root (source: "./"); no manifest change needed.

What carried over from Wisdom

Wisdom (SQL / MDLH) MCP form kept here
array_contains(tags, …) doesn't exist → join tagrelationship tags are the tags parameter, not a conditions entry
CM not in gold layer → custommetadatarelationship resolve_metadata first, then the set's display name in attributes
owner_users holds john.doe, not "John Doe" get_users first, then filter; any_conditions for user-or-group
table_columns[] GUID expansion get_assets(attributes=["columns"])
find ALL same-named instances, filter has_lineage = true same disambiguation, then traverse_lineage
recursive subdomain + product GUID collection same walk, GUIDs into domain_guids
per-DQ-tool link columns; case-sensitive pass/fail preserved verbatim — the uppercase native PASS vs lowercase partner trap
glossary QN is opaque {nanoId}@{glossaryQN} scope via glossary_qualified_name, never pattern-match the QN
mandatory asset_type filter + 63 noise types type precision, and the exclusion list only when genuinely going broad

Validation

This has not been A/B'd yet. The eval dataset and the harness that measures it are in agent-toolkit-internal (companion PR): a search_skills capability with one case per rule this skill teaches, and cli.py run --with-skills / cli.py compare to diff the arms case-by-case.

Run before merge:

make eval-ab ARGS="--capabilities search_skills"

Open questions for review

  1. Rules that duplicate the tool descriptions. Rules 1, 2, 10 and 11 (verbatim query, relax-don't-keyword, count-vs-list, pagination) already appear in the semantic_search / search_assets docstrings. If the A/B shows no movement on those cases, the right fix is to cut them from the skill, not ship them — the skill should carry what the tool surface cannot.
  2. resolve_metadata namespace values. references/filters.md uses namespace_type="classification" and "business_metadata". The latter is confirmed by the tool's own docstring; the former is inferred and unverified against the parameter's accepted values. Please confirm or correct.
  3. The 63-type noise list. Copied from gold-layer-assets.md, where it exists because gold.assets is a partitioned SQL table. Whether persona-scoped search_assets results surface those types at all is untested — if they never appear, the list is dead weight in the context window and should be trimmed.
  4. One skill or several? Twelve Wisdom skills collapsed into one skill with six references. If we want per-domain discovery to be independently loadable (e.g. a data-quality skill on its own), that is a different split.

Unrelated staleness noticed and deliberately left alone: CLAUDE.md describes 15 tools (the surface is 41), and README.md links a non-existent claude-plugin/ directory.

🤖 Generated with Claude Code

…ugins

Transposes the discovery judgment from Wisdom's `category: discover` skills
onto the MCP tool surface. Wisdom's versions are SQL over the `gold.*`
lakehouse tables; the SQL does not carry over, but the judgment does — which
tool answers which ask, and the filter rules that decide whether it returns
the right thing.

One skill, six references (progressive disclosure), replacing twelve
SQL-shaped ones. What survived the transposition:

- tags are the `tags` parameter, not a `conditions` entry
- custom metadata needs `resolve_metadata` before it can be filtered
- owner fields hold usernames, not display names — `get_users` first
- a table's columns come from `get_assets(attributes=["columns"])`
- lineage needs the right instance picked from all same-named hits
- a domain's assets need the subdomain/product walk, not just the root GUID
- DQ pass/fail values differ per source AND differ in case
- glossary qualifiedNames are opaque, so scope via glossary_qualified_name
- asset-type precision, and the internal-type exclusions when going broad

`cursor-plugin/skills` is a symlink to the repo-root `skills/`, so a skill is
authored once and ships to both plugins.

Not yet measured: the A/B run that justifies this lands with the eval dataset
in agent-toolkit-internal. Rules that duplicate what the tool descriptions
already say should be cut from the skill rather than shipped — the diff is
what tells us which those are.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@benhuds

benhuds commented Aug 19, 2026

Copy link
Copy Markdown

@ankitjaggi thanks for putting this together. I checked the references against the latest agent-toolkit-internal, then tested every finding against a live tenant. Adding a couple of comments below

(Also – your open question 2 looks fine. namespace_type="classification" works. middleware/arg_normalization.py:_resolve_enum_value maps singular to plural before validation runs. Nothing to change.)

1. lineage.md: immediate_neighbors defaults to one hop

Issue. traverse_lineage_tool has immediate_neighbors=True by default, so depth is ignored. lineage.md documents neither, and calls the page-size param size; it's limit (default 10, max 20).

Impact. A question like "what's upstream of this table?" returns immediate parents only, and nothing in the response flags that the traversal stopped, so the agent reports a partial graph as the full picture. Tested on a three-layer bronze/silver/gold pipeline: the default returned 2 upstream tables, multi-hop returned 4 and picked up the actual source layer. Both responses said has_more: false. Any question spanning more than one hop ("does this table feed the exec dashboard?", "what breaks if we drop this column?") gets a confidently wrong answer.

Proposed fix. Add to the Traverse section: one hop by default; for multi-hop, set immediate_neighbors=false and depth explicitly. Rename size to limit, max 20. Say which asks need multi-hop.


2. filters.md: directly_tagged default looks inverted

Issue. directly_tagged defaults to True (mcp_tools.py:1972, origin/main): direct tags only, propagated excluded. However, filters.md says "Default (propagated included) is usually what a governance question means."

Impact. With the default as it actually is, "which tables are tagged PII?" returns only the assets someone tagged by hand and drops everything that inherited the tag via propagation. Tested: PII returns 26 tables by default, 34 with propagation included, so about a quarter of the answer goes missing. The reference also tells the agent to state which mode it used, so it will report that inherited tags were included. The customer gets a short list labelled as the complete one, which wouldn't be good for a compliance check.

Proposed fix. Flip it: default is direct-only, pass directly_tagged=false to include propagated tags. The "be explicit about which you used" line is good and should stay – just needs the right default behind it.


Your open question 1: cut them. Rules 1, 2, 10 and 11 are near-verbatim from the semantic_search_tool docstring.

…ated rules

Review findings from #236:

- lineage.md: traverse_lineage defaults to immediate_neighbors=true, which
  makes depth a no-op and silently returns a one-hop graph with
  has_more=false. Document the default, say when to set
  immediate_neighbors=false, and use the canonical param name limit
  (default 10, max 20) instead of the legacy size alias.
- filters.md: directly_tagged defaults to true (direct tags only), the
  opposite of what the reference claimed. Flipped, with the guidance to
  pass false for governance questions that need propagated tags.
- SKILL.md: drop rules 1, 2 and 11 — near-verbatim restatements of the
  semantic_search_tool user_query description and docstring. Kept the
  count-vs-list rule, whose substance is not in the tool schema.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@ankitjaggi

Copy link
Copy Markdown
Collaborator Author

@benhuds thanks — both findings reproduce against main, all three applied in 7b90e19.

1. lineage.md / immediate_neighbors. Confirmed: immediate_neighbors: bool = True and the depth field description itself says it is ignored while that's set (mcp_tools.py:2410-2445). Rewrote the Traverse section to lead with the one-hop default, call out that has_more stays false either way, and list which asks need immediate_neighbors=false + an explicit depth.

One correction on the param name: size isn't broken — limit carries validation_alias=AliasChoices("limit", "size"), so existing calls still validate. But limit is canonical, so the doc now says limit, default 10, max 20 (clamped at mcp_tools.py:2515).

2. filters.md / directly_tagged. Confirmed inverted — directly_tagged: bool = True, "Only directly tagged (not inherited)". Flipped as proposed: default is direct-only, pass false to include propagated tags, and the be-explicit line stays.

Open question 1 — cut 1, 2, 11; kept 10. Agreed on 1, 2 and 11: they restate the user_query field description and the closing docstring sentence almost word for word, same Good/Bad example included. Kept the count-vs-list rule — return_count_only's description is only "If true, return only the total match count", which never says a limit=100 page isn't a count, and that's the actual failure mode.

Worth flagging: the A/B eval that was meant to measure these rules never ran (blocked on eval venv + proxy access), so cutting those three is a judgment call on duplication, not a measured delta. If you'd rather have the numbers before merge, I can hold on that part.

Thanks for checking question 2 on namespace_type — leaving that as is.

Analysed 30 days of the `mcp` Braintrust project, scoped to the clients this
skill ships to (claude-code + cursor-vscode): search_assets_tool 11,260 calls,
traverse_lineage_tool 5,406, semantic_search_tool 2,479.

SKILL.md
- Call-shape rule (new 10). metadata.correction_guidance over 30d shows the
  top rejects: `query` for `user_query` (53), `conditions` as a string (26),
  `filters` (24), `limit` out of range (24), `direction: "BOTH"` (10),
  `type_names`/`asset_types` (10). None of this was covered.
- Compound-ask rule (new 11) and non-catalog-ask rule (new 12): median query
  is 7 words, p90 13, and asks routinely bundle resolve + owners + tags.
- Columns rule now covers the reverse direction (which table has column X) —
  a recurring ask the forward-only rule left unanswered.
- Naming rule now lists the attributes to request: displayName appeared in 2
  of 276 calls against name's 61, and 53% of calls request no attributes at
  all, which made the rule unactionable as written.
- Popularity rule distinguishes "biggest" (rowCount) from "most used".
- Routing table: counts go to search_assets (return_count_only /
  aggregations), not semantic_search; added a data-quality row.

filters.md
- Full operator table. `within`, `endswith`, `between`, `regexp`, `fuzzy`,
  `has_any_value` were all undocumented.
- The `{"name": ..., "name2": ...}` anti-pattern: 42 calls across 11 distinct
  clients in 30 days, every one a hard error. A JSON object cannot repeat a
  key; the answer is a `within` list.
- "Missing a description" via negative_conditions + has_any_value.
- Canonical single-table resolution (name + databaseName + schemaName) and
  splitting a pasted db.schema.table instead of sending it as prose.

data-quality.md
- dqRuleTemplateName plus the observed identifiers. Traces show invented
  compounds (CUSTOM_SQL_QUERY, CUSTOM_SQL_ROW_COUNT) and UI labels
  ("Standard Deviation") passed as filter values.

lineage.md
- direction has no BOTH; 10 rejected calls in 30 days.

Sample values are synthetic; no tenant or customer identifiers included.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@ankitjaggi

Copy link
Copy Markdown
Collaborator Author

What 30 days of production traces say the skill is missing

Pulled the mcp Braintrust project for the last 30 days and scoped it to the clients this skill actually ships to. Applied in ee15eaf.

First, a correction to how I'd been framing the tool mix. The semantic_search span (36,938 calls) has client_name = null — that's Atlan's own internal search surface, not the MCP tool. The MCP tool is semantic_search_tool. For claude-code + cursor-vscode:

Tool 30d calls
search_assets_tool 11,260
traverse_lineage_tool 5,406
semantic_search_tool 2,479
get_assets_tool 2,253

search_assets is 4.5x semantic_search for our clients. The routing table implied semantic-first; it now sends counts to search_assets explicitly.

1. Call shapes — biggest measurable failure class, previously uncovered

metadata.correction_guidance for these clients on the search + lineage tools, 30d:

Reject Count
query instead of user_query (+ search_query, semantic_query, query_text) 53
conditions passed as a string, not an object 26
filters / filter — not a parameter 24
limit out of range (17 over 100, 7 under 1) 24
direction: "BOTH" 10
type_names / asset_types / type_name 10

~150 wasted calls. New rule 10 covers all of it; BOTH also went into lineage.md.

2. name2 — thanks for the nudge to look, it's real

42 calls in 30 days across 11 distinct clients (claude-code 7, cursor 7, claude-ai 4, copilot, VS Code, sheet-add-in, …), so it's a generic model failure mode rather than one bad caller:

{"name": "ad_package", "name2": "invoice_detail"}
{"name": {"operator": "contains", "value": "EVENT"}, "name2": {"operator": "contains", "value": "SCAN"}}

The intent is always "one field, several values", and a JSON object can't repeat a key. Every one of these hard-errors — suggest_fields in utils/search.py:383 even strips the trailing digit and replies "did you mean name". The right shape is a within list, which was undocumented, along with endswith, between, regexp, fuzzy and has_any_value. filters.md now has the full operator table and names this anti-pattern directly.

3. Other gaps

  • Reverse column lookup. Rule 7 only went table → columns. "which table has a customer_email column" is a recurring ask with no guidance.
  • Rule 4 was unactionable. displayName was requested in 2 of 276 calls against name's 61, and 53% of calls request no attributes at all. You can't label by a field you didn't fetch — the rule now carries the attribute list.
  • DQ template vocabulary. Traces invent CUSTOM_SQL_QUERY / CUSTOM_SQL_ROW_COUNT and pass UI labels ("Standard Deviation") as filter values. dqRuleTemplateName was missing from data-quality.md entirely, though the link field you flagged earlier was already correct.
  • Compound asks. Median query is 7 words, p90 13, and they bundle resolve + owners + tags — none of which is one condition.
  • Pasted db.schema.table goes to semantic_search as prose instead of splitting into name + databaseName + schemaName.

Caveat on the lineage default from my last comment

I said one-hop was silently truncating. The trace evidence is weaker than that for our clients. In the 30–18 day window 80% of lineage calls ran one-hop (True or default), but the most recent window is 86% immediate_neighbors=false — and 334 of 392 payloads there are byte-identical (depth=4, limit=20), i.e. one automated caller, not organic behaviour. No commit changed the default or its description in that period. Documenting it is still right; "confidently wrong answers everywhere" overstated it.

Two things I checked and dropped rather than ship

  • snake_case condition keys (certificate_status, owner_users) looked like a gap. They're explicitly supported (mcp_tools.py:1920) and filters.md already said so.
  • The DQ link field dqRuleBaseDatasetQualifiedName was already documented correctly.

All sample values in the diff are synthetic. I also genericized a pre-existing real-company table name in the naming example.

Still no A/B eval numbers behind any of this — it's trace evidence, not a measured delta. If you want the eval run before merge, the rules most worth measuring are 10 and the within addition.

@ankitjaggi

Copy link
Copy Markdown
Collaborator Author

A/B eval results — the numbers benhuds asked for

Ran the golden set twice against ai-eval.atlan.com, with and without the skill loaded, using the harness in atlanhq/agent-toolkit-internal#458. Two datasets: the original 14 search_skills cases, and 14 new search_traces cases mined from 30 days of production traces (8724b47a, rubrics corrected in 830135af).

The new trace-derived cases: 4/14 → 9/14, no regressions

                before   after       Δ
ArgsCheck          62%    100%   +38 pts
Selection          62%    100%   +38 pts
OrderFollowed       0%    100%  +100 pts
NoToolErrors      100%    100%    +0
LLMJudge           69%     69%    +0
NoWrongTool         50%     50%    +0

Fixed: tr_multi_value_one_field, tr_reverse_column_lookup, tr_dq_rules_for_table, tr_missing_description, tr_biggest_not_most_popular. compare exits 0 — no regressed case.

Across three independent baseline runs the baseline scored 6, 6, 4 of 14; the skills arm scored 10, 10, 9. Individual case attribution is noisy at n=14, but the aggregate separation is consistent.

The original cases: 7/14 → 8/14

ArgsCheck   86% → 100%     Selection  86% → 100%
LLMJudge    64% →  64%     OrderFollowed 100% → 67%

This gap is the most useful thing in the run. The original dataset was written to guard the rules the skill had before review — including the four you identified as near-verbatim from the semantic_search_tool docstring. It moves +1. The trace-derived cases move +5. That is independent evidence that cutting rules 1, 2 and 11 was correct: the tool descriptions were already carrying them, so guarding them in the skill bought nothing.

Its one "regression" is a case-design bug, not a skill regression. sk_owner_resolve_first: with the skill loaded the agent called get_users_tool, found no user named "John Doe", and stopped. LLMJudge passed. It failed only OrderFollowed, which mandates a follow-up search_assets call — while the case's own expect explicitly allows "says clearly that no such user exists". Baseline made four calls including a wasted semantic_search first and got the credit. The skills arm was correct and cheaper, and the rubric punished it.

Honest limits

LLMJudge is flat, in both datasets. The skill decisively improves how the agent calls the tools — right tool, right arguments, right order — and shows no measurable improvement in final answer quality as scored by a blind judge. I am not going to dress that up: on this tenant, for these 28 asks, the mechanical gains did not translate into answers the judge rated higher. Worth knowing before anyone claims the skill makes answers better.

I burned two runs on my own bad rubrics. My first pass wrote mechanism into expect ("filtered on the template identifier", "read off the asset"). The judge never sees tool calls, so it marked correct answers wrong — tr_business_names_requested scored 0.4 for correctly reporting that no displayName is set. Two A/B runs each reported a bogus REGRESSION from this. Fixed in 830135af: every expect is answer-only, mechanism asserted via args / tool_order / NoToolErrors.

Rule 12 still fails. tr_not_a_catalog_question fails in both arms — on "we're losing our Atlan license, back up all our metadata" the agent called 6-12 tools instead of abstaining. The abstain rule I added does not work. @ankitjaggi has said to leave it for now; flagging so it is not mistaken for a passing rule.

Pre-existing cases worth a separate look (not touched, to avoid retuning rubrics mid-measurement): sk_dq_failing_multisource demands four DQ sources when the tenant has three; sk_cm_resolve_first and sk_columns_from_table sit at 0.75 against a 0.8 threshold on rubric specificity rather than substance.

Infrastructure note

Concurrency 16 saturates the inference proxy — connect errors, agent and judge timeouts, and judge scores coming back as "failed to parse". Every run at concurrency 4 completed with zero infrastructure errors. The A/B in the Makefile should probably drop its default concurrency for the skills arm, since a 27k-char system prompt on every case is what tips it over.

@abhinavmathur-atlan

Copy link
Copy Markdown
Collaborator

@Aryamanz29 to complement these lets also add a line in server instructions on using these skills in case the user has them?

New rule 7, from a customer evaluation where the same scenario passed with a
real business-unit value and failed on all three models with a plausible but
non-existent one.

Rule 6 already says tags and custom metadata are not `conditions` — the slot.
Nothing covered the value. A wrong value fails silently: the search returns an
empty set, which the agent reports as "nothing is tagged that way". In the
customer's traces the models cycled through eight invented business-unit
values across three models and never once called resolve_metadata to list the
real ones. The same scenario with a correct value passed 3/3.

The rule: resolve_metadata first, filter on a value from that list, and if
nothing matches report which values do exist rather than answering zero. Plus
the reason it matters — "no PII assets in that business unit" is a dangerous
sentence when the truth is that the unit name was misspelled.

Example values are synthetic; renumbers former rules 7-12 to 8-13.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@ankitjaggi

Copy link
Copy Markdown
Collaborator Author

Customer evaluation traces — one new rule, and two product bugs a skill can't fix

A customer ran 19 scenarios against the MCP across Haiku/Sonnet/Opus (11/19, 14/19, 15/19). I pulled their tenant's traces for the last 30 days — 1,157 spans — and read the actual tool calls behind each failure. Tenant and asset names stay out of this repo; the example values in the diff are synthetic.

New rule 7: resolve tag and custom-metadata values before filtering (1a3d50a)

The clearest signal in the whole set is a pair of near-identical scenarios:

  • "assets tagged 'contains PI' owned by Business Unit A" — failed on all three models
  • "assets tagged 'contains PI' owned by Business Unit B" — passed on all three

Same shape, same slots. The only difference is that B is a real business-unit value and A is not. The models correctly worked out that business unit is an Atlan tag rather than an attribute — rule 6's job — and then guessed the value eight different ways across the three runs, never once calling resolve_metadata to enumerate the real ones.

Rule 6 covered the slot. Nothing covered the value, and a wrong value fails silently — an empty result that gets reported as "nothing is tagged that way". On a PI compliance question that is the worst possible failure mode, which is why the rule ends with: "no PII assets in that business unit" is a dangerous thing to say when the real answer is that you spelled the unit wrong.

Also confirmed in that same failing scenario, and already fixed on this branch:

  • atlanTags2 / atlanTags3 / atlanTags4 inside any_conditions — the numbered-key anti-pattern, on the tag field, doubly broken because atlanTags isn't searchable as a condition at all. This is what within is for.
  • directly_tagged: true on nearly every call (the default) — one call out of the whole set used false. Exactly the propagation default you flagged, on a compliance question.
  • conditions passed as a JSON string; {"__typeName": {"operator": "contains", "value": "Table"}} instead of asset_type; invented fields assetTags, businessOwner.

Two product bugs — please route these, they are not skill-fixable

1. Lineage truncates silently. For the "show complete upstream datasets" scenario the model did everything right:

{"direction": "UPSTREAM", "depth": 5, "limit": 20, "immediate_neighbors": false}

The response contains nodes with traversalOrder up to 50 while limit caps at 20. I grepped all 67 of this tenant's lineage responses for has_more, truncated, exceeded, Load more — zero hits. The graph is cut to 20 nodes and nothing says so, so every model reports a partial graph as complete. Failed 3/3.

2. Lineage is flooded with per-DAG-run artifacts. The upstream set is dominated by GCSObject entries with isPartial: true, one per execution date (…_20250324T000000_dev_aarch64, …_20260819T000000_…, …_20260820T090000_…). With a 20-node budget, transient scratch-bucket objects crowd out the actual upstream datasets the user asked for.

Together these produce a genuinely counterintuitive result: going deeper makes the answer worse. On the shallower "initial upstream" variant of the same question, Haiku passed with immediate_neighbors: true (clean direct parents) while Sonnet and Opus failed at depth 10–20, having drowned in run artifacts. Depth is currently a liability on this tenant.

Correction to my earlier comment

I previously attributed this customer's lineage failure to the immediate_neighbors one-hop default. The traces don't support that. 51 of their 67 lineage calls passed immediate_neighbors: false with real depths, including the failing scenario. The default trap is real — 16 of 67 calls hit it, 13 of them passing depth: 20 while immediate_neighbors stayed true, so the depth was silently ignored — but it is not what broke this scenario. Documenting the default remains right; it just isn't this bug.

What worked

10 of 19 passed on all three models, and they share one shape: an exact identifier in, one tool call out — owner lookup, full path, column list, sample rows (6–13s), a count via return_count_only (13s on all three), single-asset PI status, a targeted description update. The tool surface is solid for identifier-driven questions. Every failure in the set is either a value the agent had to discover, or a graph the tool truncated without saying so.

One data-quality note on the customer's table

Rows 7 and 8 don't match their listed prompts — the "top users of X in last 30 days" summary belongs to prompt 8, and prompt 7 ("if dataset freshness_metrics is deprecated, which downstream assets are impacted?") has no row at all. Worth resolving before anyone reads those two verdicts, since one of them is a lineage-impact question.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants