fix(converter): keep <br> as a hard break in html → markdown lists and tables - #271
Merged
Merged
Conversation
…d tables html → markdown turned <br/> in list items and table cells into a space, while storage → markdown has kept them since #248 (a backslash hard break in list items, an inline <br> in cells), so the converters disagreed on the same markup. Move the walker's hard-break helpers (collapseInline, escapeContinuationLine and the trailing-backslash handling, now resolveHardBreaks) into markdown-cleanup so both converters share them. In html → markdown, mark each <br> with HARD_BREAK outside code spans, resolve it per list item, and write it as an inline <br> in table cells. Backslash and block-opener character references are decoded where they decide how a break is escaped, since entities are otherwise decoded after list items are rendered. Top-level <br> stays a soft break. Fixes #253
html → markdown decodes entities only after list items are rendered, so the break handling looked at undecoded text. A continuation line that only decodes to a block opener (`- b`, `----`, `&#35; b`) was not escaped, a reference to a sentinel codepoint could become a live sentinel, and ` ` next to a break left a stray backslash. Extract the converter's entity pass into one decodeEntities function and let resolveHardBreaks take decode/encode hooks: the opener and trailing-backslash checks look at the decoded line, a line that needs escaping is written back with & and < re-encoded, and other lines are kept as written. Drop lines next to a break before the collapse, only when a break is present. Treat <BR> as a break, as HTML tag names are case-insensitive.
|
🎉 This PR is included in version 2.27.3 🎉 The release is available on: Your semantic-release bot 📦🚀 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
After #248, storage → markdown keeps
<br/>inside list items as a backslash hard break and inside table cells as an inline<br>. html → markdown (lib/html-to-markdown.js, used byconvert --input-format html) still turned both into a space, so the two converters disagreed on the same markup (the kind of drift #149 was about).Changes
lib/markdown-cleanup.jsso both converters share them:collapseInline,escapeContinuationLine, andresolveHardBreaks(joins lines with a backslash break and doubles an odd run of trailing backslashes).StorageWalkernow calls them; its behavior is unchanged and all of its existing tests pass untouched.<br>is marked with the existingHARD_BREAKsentinel (outside code spans, where a break cannot live and stays a space), whitespace is collapsed around it, and it is resolved per item. Continuation lines are indented to the item's content column by the existing list rendering, and lines that would start a block (-,1.,#,>,---, fences, table delimiter rows) are escaped.<br>. The final tag-stripping pass would remove a literal<br>, so the cell text carries<br>, which the entity pass turns back into<br>.decodeEntitiesfunction, andresolveHardBreakstakesdecode/encodehooks: the block-opener and trailing-backslash checks look at the decoded line (so- b,----,&#35; bor a backslash written as\are handled), a line that needs escaping is written back with&and<re-encoded (the passes before the final decode would otherwise treat them as the start of a reference or a tag), and every other line is kept exactly as written. lines next to a break are dropped before the collapse, and only when a break is present, so text without breaks is untouched.<BR>counts as a break in lists and cells, since HTML tag names are case-insensitive (the walker's XML parser is case-sensitive and does not).<br/>is still a soft break in both converters (out of scope, as noted in the issue).Behavior notes
<br>inside a list item or table cell produces different output; everything else is unchanged.Testing
10.), nesting, whitespace and consecutive breaks, edge breaks, trailing backslashes (literal and as character references), the full set of block-start lines (including ones written as character references), a literal fence line, code spans, table cells, and a literal U+E003.<br/>tests, plus character-reference and variants, through bothhtmlToMarkdownandstorageToMarkdownand expect identical output.main.npm test(40 suites, 1730 tests) andeslintpass.Not covered here: inline
<code>containing backticks, which html → markdown wraps in single backticks regardless of content (an existing limitation unrelated to breaks).Fixes #253