Repository navigation
Model comments as their own node type rather than raw HTML #36
Description
Activity
MariusStorhaug commented
on Aug 15, 2026 MemberAuthorMore actionsCorrected the unterminated-comment rule. The description said an unterminated comment runs to the end of the document; §4.6 says it ends at
the first subsequent line that meets a matching end condition, or the last line of the document, **or the last line of the container block containing the current HTML block**.So a stray
<!--inside a block quote or a list item swallows the rest of that container and stops there, rather than the rest of the file. Only a comment at document level runs to the end of the document. Two places in the implementation plan updated to match, including a test for the container case which the original wording would have missed.Found by the session writing the specification, checking the rule against the spec text rather than against this issue.
MariusStorhaug commented
on Aug 15, 2026 MemberAuthorMore actionsTwo things for this issue, both settled in #37.
The open one-type-or-two decision is resolved as two, and the reasoning is recorded in
docs/markdown-object-model/design.mdunderAlternatives considered. A PowerShell class has a single base, and comments occur at block and inline level, so one type cannot derive from bothMarkdownBlockandMarkdownInline. The model carriesMarkdownCommentBlockandMarkdownComment, mirroring theMarkdownHtmlBlockandMarkdownRawHtmlpair already in the inventory, so$_ -is [MarkdownBlock]stays a complete filter. Both report aTypeofComment, soDescendants('Comment')finds every comment at either level. The other two options are recorded as rejected: a single type deriving from the node base, which sits outside the block and inline split the filtering idiom depends on, and anIsCommentflag on the existing HTML nodes, which leaves the content as raw text with its delimiters attached.The
Open:heading in the description is therefore stale, and the two specification tasks in the implementation plan are delivered by that pull request.The cited CommonMark example number is wrong. The
<!-- foo -->*bar*case is example 177, not 172 — example 172 is the<style type="text/css">case. Verified by counting example fences inspec.txtfor 0.31.2, which has 652 of them. The point the example is cited for is unaffected: the HTML block ends at the first line containing-->and the trailing*bar*belongs to the same block, so a comment with content after it on the same line stays an HTML block. The specification's own note beside it reads "anything on the last line after the end tag will be included in the HTML block".Everything else in this issue's description matched what went into the specification and the design.
MariusStorhaug commented
on Aug 15, 2026 MemberAuthorMore actionsCorrected the CommonMark example reference: the
<!-- foo -->*bar*case is example 177, not 172. Example 172 is a<style>block.Verified against the published example set rather than by counting:
spec.jsonfor 0.31.2 contains 652 examples, and the one whose markdown is<!-- foo -->*bar*is number 177 in the HTML blocks section.Worth recording how the error survived. I originally wrote 172 from memory of the spec text. Recounting against
spec.txton thecommonmark-specmaster branch gave 179 \u2014 also wrong, because master has drifted from the published 0.31.2 release. Only the publishedspec.jsonfor the exact version gives the number that matches the anchor a reader will follow.Found by the session writing the specification, which checked the number rather than copying it from this issue.
Comments are how a Markdown document carries instructions that are not content. Linter directives (
<!-- markdownlint-disable MD041 -->), formatter pragmas (<!-- prettier-ignore -->), generated-region markers (<!-- BEGIN GENERATED -->…<!-- END GENERATED -->), and editorial notes to the next author all travel as comments, because Markdown has no other way to say something to a tool rather than to a reader.Request
Current experience
Markdown has no comment syntax of its own. A comment is HTML, so under the object model a comment arrives as raw HTML — a block-level one as an HTML block, an inline one as raw HTML — carrying its delimiters as part of an opaque string. Nothing distinguishes it from a
<div>or a<details>.Everything a caller wants to do with a comment therefore starts with string matching. Finding the generated region means scanning raw HTML nodes for text beginning
<!--and ending-->, then slicing the delimiters off to read what is inside. That is the same pattern matching against raw text the object model exists to remove, moved one layer in and made less obvious.The consequences show up in the capabilities built on the model:
<!-- markdownlint-disable MD041 -->, so the failure is not hypothetical.Desired experience
A comment is a thing the model knows about. It can be found, read, written, and removed without touching delimiters or matching text.
Acceptance criteria
<!--and-->delimiters, and exposes whether it was terminated in the source.Out of scope
Technical decisions
This has to be settled in 1.3, not added later. Adding a node type is normally additive and non-breaking, but this one is not: it changes what the parser returns for input that already parses. A caller matching
Type -eq 'HtmlBlock'in 1.3 would silently stop matching comments when a comment type arrived in 1.4. That is exactly the rule #8 sets for itself — anything that changes the shape of whatConvertFrom-Markdownreturns is a 1.3 decision. Nothing has shipped yet; the latest release is v1.2.5.Open: one type or two. Comments occur at both block and inline level, and PowerShell classes have single inheritance, so a single type cannot derive from both
MarkdownBlockandMarkdownInline. Three options:MarkdownHtmlBlock/MarkdownRawHtmlsplit already in #8, so$_ -is [MarkdownBlock]keeps working. Two names for one concept, andDescendants('Comment')needs both to report the sameType.IsCommentflag on the existing HTML nodesThe existing split is the strongest precedent, but this is the decision that shapes the rest and it is not settled here.
A block-level comment is only a comment when the block is exactly a comment. CommonMark ends an HTML block at the first line containing
-->, and everything after the terminator on that line belongs to the same block — example 177 shows<!-- foo -->*bar*producing one HTML block in which*bar*is not emphasized. So a block that is a comment plus trailing content stays an HTML block; promoting it would silently drop the trailing text.Unterminated comments run to the end of their container. The end condition is a line containing
-->, or the last line of the document, or the last line of the container block containing the comment (§4.6). A stray<!--at document level therefore swallows the rest of the file; inside a block quote or list item it swallows the rest of that container and no more. Both are real authoring accidents and ones a validator should be able to report. The model records whether the comment was terminated rather than pretending it was.The degenerate forms are comments. §6.6 admits
<!-->and<!--->as HTML comments alongside the usual form. They carry no text and must round-trip as written.Delimiters and inner spacing are preserved. Following the round-trip contract in #8, the node keeps enough of the source form to re-render it as written —
<!-- x -->does not become<!--x-->. Inner text is exposed separately for reading.Not a dialect concern. Comments are core CommonMark at both block and inline level, so this belongs with the base grammar and not with #30.
Implementation plan
Specification
docs/markdown-object-model/spec.mddocs/markdown-object-model/design.mdModel
Parser
<!-->and<!--->formsRenderer
Tests
Documentation