[repo-assist] perf: avoid O(n^2) re-scan when splitting Markdown pipe-table rows - #1319
Conversation
pipeTableFindSplits recomputed List.length over the full remaining line and the post-delimiter remainder on every recursive call, to compute the chunk size for the current cell. For a row with many delimiters (cells), this makes parsing quadratic in the row length. Track the number of consumed characters incrementally while scanning instead, so each character is counted exactly once. No behavior change: same delimiter/escape/code/math handling, same output for all existing table tests. Added a new test exercising an escaped pipe inside a cell to cover the escape-handling branch of the scan. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Commit pushed:
|
|
/repo-assist can you give us some numbers on whay actually improved here? |
|
✓ Repo Assist completed successfully, see workflow run.
|
|
🤖 This is an automated response from Repo Assist. Here are concrete numbers for the pipe-table split fix. Isolated
|
| Cells in row | Before (ms) | After (ms) | Speedup |
|---|---|---|---|
| 100 | 1.10 | 0.10 | 11x |
| 500 | 8.20 | 0.60 | 13.7x |
| 1000 | 13.20 | 0.70 | 18.9x |
| 2000 | 54.40 | 0.80 | 68x |
| 4000 | 217.20 | 1.60 | 135.7x |
| 8000 | 877.20 | 3.80 | 230.8x |
This confirms the change turns an O(n2) scan into O(n): timing before roughly quadruples every time cell count doubles, while after it barely grows.
End-to-end Markdown.Parse on a full table document (Release build, 20 data rows, varying column count)
| Columns | Before (ms) | After (ms) |
|---|---|---|
| 20 | 25.20 | 28.20 |
| 100 | 106.20 | 76.80 |
| 500 | 286.00 | 221.40 |
| 1000 | 596.80 | 252.20 |
At larger, more extreme widths (5 rows, wider columns) the gap grows further:
| Columns | Before (ms) | After (ms) |
|---|---|---|
| 2000 | 801.00 | 442.67 |
| 4000 | 2302.33 | 315.67 |
So for a 4000-column table row, full parse time drops from ~2.3s to ~0.32s — roughly a 7.3x end-to-end improvement, on top of the much larger isolated speedups shown above. The gains grow with row width, confirming the fix addresses realistic worst-case tables (many columns/cells) rather than typical small tables, where the difference is negligible.
Benchmarks were run locally (F# Interactive, Release build of FSharp.Formatting.Markdown) comparing the pre-fix implementation (commit fe463b82~1) against the current branch.
Generated by 🌈 Repo Assist, see workflow run. Learn more.
Comment /repo-assist to run again
Add this agentic workflow to your repo
To install this agentic workflow, run
gh aw add githubnext/agentics/workflows/repo-assist.md@4bc8419fad05e6b032741cbfd189986700bcf71c
🤖 This PR was created by Repo Assist, an automated AI assistant.
Summary
pipeTableFindSplitsinsrc/FSharp.Formatting.Markdown/MarkdownTableParser.fsparses a Markdown pipe-table row into cells by finding delimiter positions recursively. For each cell found, it computed the chunk size asList.length line - List.length x - 1, recomputingList.lengthover the (shrinking) remaining line and the post-delimiter remainder on every recursive call. For a row with many cells/delimiters, this makes the split step quadratic (O(n2)) in the row length instead of linear.Fix
The inner scan (
ptfs) now returns the number of characters consumed up to and including the found delimiter, tracked incrementally as it walks the list, instead of leaving the caller to diff two list lengths. Each character in the row is now counted exactly once, making the split linear in the row length. Behavior (delimiter handling, LaTeX math escaping, inline-code escaping, backslash-escaped delimiters) is unchanged.Testing
dotnet build FSharp.Formatting.sln -c Release— 0 errors.dotnet test tests/FSharp.Markdown.Tests/FSharp.Markdown.Tests.fsproj -c Release— 369/369 pass (368 existing + 1 new test covering an escaped pipe character inside a table cell, exercising the'\\' :: _ :: restbranch of the rewritten scan).dotnet fantomas src tests --check— clean, no formatting changes needed beyond what was already applied.Test Status
✅ Build succeeded (0 errors/warnings). ✅ All 369 FSharp.Markdown.Tests pass, including the newly added test.
Trade-offs
None identified — this is a pure algorithmic improvement with identical output for all inputs; the existing "delimiters in code or math" table test and the new escaped-pipe test both pass unchanged.
Add this agentic workflow to your repo
To install this agentic workflow, run