Skip to content

fix(converter): escape literal fence-like prose so it cannot capture code blocks - #269

Merged
pchuri merged 2 commits into
mainfrom
fix/literal-backtick-fence
Oct 4, 2026
Merged

pchuri merged 2 commits into
mainfrom
fix/literal-backtick-fence

Conversation

@pchuri

@pchuri pchuri commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Problem

A prose line that starts with three or more backticks is indistinguishable from a fence the converter emitted itself. splitOnFences pairs it with the next real fence, so the real code body falls outside the fence and goes through whitespace cleanup. The code is changed silently.

htmlToMarkdown('<ul><li>a<ul><li>b</li></ul>```</li></ul><pre><code>  indented\n    code   here</code></pre>')
// the <pre> body loses its indentation

This affects both converters. Besides html → markdown (convert --input-format html, reported in #254), storage → markdown is affected too, which is the read path: a <p>```</p> followed by a code macro loses the indentation of the code body.

Changes

  • New shared helper escapeFenceLikeText() in lib/markdown-cleanup.js. It backslash-escapes every backtick of a line-leading run of 3+ backticks in prose (``` becomes \`\`\`). Escaping only the first would leave the rest free to pair with later backticks into a code span on write-back. NBSP (literal or as an entity) counts as indentation before the run.
  • storage → markdown: applied in StorageWalker.renderText(), except inside code spans and link labels (which already escape backticks).
  • html → markdown: applied before conversion to the prose between tags. <pre> and <code> regions and tag attributes are left alone. Backtick entities (&#96;, &#x60;) in prose are decoded first, because the entity pass runs after the fence-aware cleanup and would otherwise bring back a literal fence line. Only < followed by a letter, / or ! is treated as a tag (so prose like a < b is not skipped), and the <pre>/<code> region patterns are used only when a closing tag exists, which avoids a quadratic rescan on unclosed <pre>.
  • Code bodies, code spans, attribute values and mid-line backtick runs are unchanged.

Output change

A literal fence-like line in prose is now written as \`\`\` instead of ```. The old output turned into a real code block when written back to Confluence; the escaped form round-trips as prose. Five existing expectations that pinned the unescaped output (#243, #244) were updated to the escaped form. One of them has loose text directly after a callout paragraph with no blank line, which Markdown treats as a continuation of that paragraph, so its write-back assertion now expects that.

Testing

  • Unit tests for the helper, including that an escaped line cannot open a fence.
  • html → markdown and storage → markdown regression tests: the repro from Literal ``` after a nested list corrupts a later code block in html → markdown #254, a top-level paragraph, <br> before the literal, backtick entities and CDATA, a callout, and round-trip. Also checks that code bodies, code spans, mid-line runs and attribute values are untouched.
  • npm test (40 suites, 1578 tests) and eslint pass.

Fixes #254

pchuri added 2 commits October 4, 2026 21:01
…code blocks

A prose line starting with 3+ backticks was indistinguishable from a
fence the converter emitted. splitOnFences paired it with the next real
fence, so the real code body fell outside the fence and went through
whitespace cleanup, silently changing the code. This affected both
html → markdown and storage → markdown (the read path).

Escape the first backtick of a line-leading run in prose
(``` becomes \```) in both converters, through one shared helper.
Code bodies, code spans, attribute values and mid-line runs are left
alone, and backtick entities in html prose are decoded first so they
cannot reintroduce a literal fence. The escaped text also round-trips
as prose instead of turning into a code block on write-back.

Fixes #254
- Escape every backtick of the leading run, not just the first, so the
  rest cannot pair with later backticks into a code span on write-back.
- Treat NBSP (literal, &nbsp;, &#160;, &#xA0;) as indentation before the
  run, in both converters.
- Only treat `<` followed by a letter, `/` or `!` as a tag in
  html → markdown, so prose like `a < b` no longer hides a fence line.
- Skip the <pre>/<code> region patterns when no closing tag exists, which
  removes a quadratic rescan on unclosed <pre>.
- Document that text starting at an inline-element boundary counts as a
  line start.
@pchuri
pchuri merged commit 4ad48ce into main Oct 4, 2026
6 checks passed
@pchuri
pchuri deleted the fix/literal-backtick-fence branch October 4, 2026 12:19
github-actions Bot pushed a commit that referenced this pull request Oct 4, 2026
## [2.27.1](v2.27.0...v2.27.1) (2026-10-04)

### Bug Fixes

* **converter:** escape literal fence-like prose so it cannot capture code blocks ([#269](#269)) ([4ad48ce](4ad48ce)), closes [#254](#254)
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

🎉 This PR is included in version 2.27.1 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Literal ``` after a nested list corrupts a later code block in html → markdown

1 participant