Skip to content

Leverage execution policies in dbShapes and PLC decomposition - #2367

Open
nikosavola wants to merge 2 commits into
KLayout:parallelizefrom
nikosavola:nikosavola/push-olzlkmqortox
Open

nikosavola wants to merge 2 commits into
KLayout:parallelizefrom
nikosavola:nikosavola/push-olzlkmqortox

Conversation

@nikosavola

@nikosavola nikosavola commented Jun 3, 2026 •

Copy link
Copy Markdown
Contributor

Same idea as #2364, but for dbShapes and the PLC convex decomposition.

  • dbShapes.cc: layer_op::erase/redo sorts the shape vector before the erase walk.
  • dbPLCConvexDecomposition.cc: Hertel-Mehlhorn sorts the newly inserted points and, for each vertex, the edges around that vertex by angle and length.

Both sorts now use std::execution::par where the standard library supports execution policies, and stay plain std::sort where it doesn't (macOS/libc++, MSVC without TBB, C++11 builds). The guards are __cpp_lib_execution and __has_include(<execution>).

Benchmark

Two cases from the db_parallel_benchmarks Google Benchmark suite on the parallelize-rebased branch of my fork, on the same machine and with the same back-to-back serial/parallel procedure as #2365 (Xeon Gold 6248 @ 2.5 GHz, 16 cores, GCC 13.3, -O2, TBB). shapes_erase fills a flat Shapes container and erases every second polygon, which is mostly sorting. plc_decomposition runs convex decompositions of a small concave polygon (Linux only).

shapes_erase, 16384 polygons:

CPUs serial (ms) par (ms) speedup
1 5.042 4.662 1.08x
2 5.244 4.506 1.16x
4 5.246 4.276 1.23x
8 5.270 4.132 1.28x
16 5.249 4.313 1.22x

Speedup versus number of CPUs

  • The erase benefits the most: up to 1.28x at 8 cores, since the operation is nearly all sort. At 4096 polygons the gain is smaller (about 1.09x from 2 cores on) and turns slightly negative at 16 cores (0.97x), which is the parallel sort overhead on a short array.
  • The PLC sorts time the same in both builds (0.97x to 1.01x at every CPU count). Hertel-Mehlhorn sorts very small vectors (the points of one insertion step, the edges around one vertex), so there is little to parallelize. I kept that hunk for consistency with the other sorts, but it is easy to drop or guard by size if you prefer.
  • The Apple M5 run showed the same behaviour, so nothing here is x86-specific.

@nikosavola
nikosavola changed the base branch from master to parallelize June 22, 2026 14:00
@nikosavola
nikosavola marked this pull request as ready for review September 25, 2026 07:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant