Skip to content

[repo-assist] perf: avoid O(n^2) re-scan when splitting Markdown pipe-table rows - #1319

Merged
nojaf merged 6 commits into
mainfrom
repo-assist/perf-pipe-table-split-20260911-9c64efc07655fd6d
Sep 28, 2026
Merged

nojaf merged 6 commits into
mainfrom
repo-assist/perf-pipe-table-split-20260911-9c64efc07655fd6d

Conversation

@github-actions

Copy link
Copy Markdown
Contributor

🤖 This PR was created by Repo Assist, an automated AI assistant.

Summary

pipeTableFindSplits in src/FSharp.Formatting.Markdown/MarkdownTableParser.fs parses a Markdown pipe-table row into cells by finding delimiter positions recursively. For each cell found, it computed the chunk size as List.length line - List.length x - 1, recomputing List.length over the (shrinking) remaining line and the post-delimiter remainder on every recursive call. For a row with many cells/delimiters, this makes the split step quadratic (O(n2)) in the row length instead of linear.

Fix

The inner scan (ptfs) now returns the number of characters consumed up to and including the found delimiter, tracked incrementally as it walks the list, instead of leaving the caller to diff two list lengths. Each character in the row is now counted exactly once, making the split linear in the row length. Behavior (delimiter handling, LaTeX math escaping, inline-code escaping, backslash-escaped delimiters) is unchanged.

Testing

  • dotnet build FSharp.Formatting.sln -c Release — 0 errors.
  • dotnet test tests/FSharp.Markdown.Tests/FSharp.Markdown.Tests.fsproj -c Release — 369/369 pass (368 existing + 1 new test covering an escaped pipe character inside a table cell, exercising the '\\' :: _ :: rest branch of the rewritten scan).
  • dotnet fantomas src tests --check — clean, no formatting changes needed beyond what was already applied.

Test Status

✅ Build succeeded (0 errors/warnings). ✅ All 369 FSharp.Markdown.Tests pass, including the newly added test.

Trade-offs

None identified — this is a pure algorithmic improvement with identical output for all inputs; the existing "delimiters in code or math" table test and the new escaped-pipe test both pass unchanged.

Generated by 🌈 Repo Assist, see workflow run. Learn more.
Comment /repo-assist to run again

Add this agentic workflow to your repo

To install this agentic workflow, run

gh aw add githubnext/agentics/workflows/repo-assist.md@ae8d551f07c7ed7619f8c58c7bb4c3ac89395d38

pipeTableFindSplits recomputed List.length over the full remaining line
and the post-delimiter remainder on every recursive call, to compute
the chunk size for the current cell. For a row with many delimiters
(cells), this makes parsing quadratic in the row length.

Track the number of consumed characters incrementally while scanning
instead, so each character is counted exactly once. No behavior change:
same delimiter/escape/code/math handling, same output for all existing
table tests. Added a new test exercising an escaped pipe inside a cell
to cover the escape-handling branch of the scan.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@dsyme
dsyme marked this pull request as ready for review September 14, 2026 12:54
dsyme and others added 2 commits September 15, 2026 21:08
@github-actions

Copy link
Copy Markdown
Contributor Author

Commit pushed: dc8971c

Generated by 🌈 Repo Assist, see workflow run. Learn more.

@nojaf

nojaf commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

/repo-assist can you give us some numbers on whay actually improved here?

@nojaf nojaf closed this Sep 28, 2026
@nojaf nojaf reopened this Sep 28, 2026
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

✓ Repo Assist completed successfully, see workflow run.

Generated by 🌈 Repo Assist, see workflow run. Learn more.

@github-actions

Copy link
Copy Markdown
Contributor Author

🤖 This is an automated response from Repo Assist.

Here are concrete numbers for the pipe-table split fix.

Isolated pipeTableFindSplits benchmark (in-memory, JIT-warmed, avg of multiple reps)

Cells in row Before (ms) After (ms) Speedup
100 1.10 0.10 11x
500 8.20 0.60 13.7x
1000 13.20 0.70 18.9x
2000 54.40 0.80 68x
4000 217.20 1.60 135.7x
8000 877.20 3.80 230.8x

This confirms the change turns an O(n2) scan into O(n): timing before roughly quadruples every time cell count doubles, while after it barely grows.

End-to-end Markdown.Parse on a full table document (Release build, 20 data rows, varying column count)

Columns Before (ms) After (ms)
20 25.20 28.20
100 106.20 76.80
500 286.00 221.40
1000 596.80 252.20

At larger, more extreme widths (5 rows, wider columns) the gap grows further:

Columns Before (ms) After (ms)
2000 801.00 442.67
4000 2302.33 315.67

So for a 4000-column table row, full parse time drops from ~2.3s to ~0.32s — roughly a 7.3x end-to-end improvement, on top of the much larger isolated speedups shown above. The gains grow with row width, confirming the fix addresses realistic worst-case tables (many columns/cells) rather than typical small tables, where the difference is negligible.

Benchmarks were run locally (F# Interactive, Release build of FSharp.Formatting.Markdown) comparing the pre-fix implementation (commit fe463b82~1) against the current branch.

Generated by 🌈 Repo Assist, see workflow run. Learn more.
Comment /repo-assist to run again

Add this agentic workflow to your repo

To install this agentic workflow, run

gh aw add githubnext/agentics/workflows/repo-assist.md@4bc8419fad05e6b032741cbfd189986700bcf71c

@nojaf
nojaf merged commit 8a05e38 into main Sep 28, 2026
4 checks passed
@nojaf
nojaf deleted the repo-assist/perf-pipe-table-split-20260911-9c64efc07655fd6d branch September 28, 2026 11:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants