diff --git a/.agents/skills/testing-sysml-repl/SKILL.md b/.agents/skills/testing-sysml-repl/SKILL.md index 62daf15a23..93390a4dfb 100644 --- a/.agents/skills/testing-sysml-repl/SKILL.md +++ b/.agents/skills/testing-sysml-repl/SKILL.md @@ -2530,9 +2530,66 @@ its output rather than in an exit code — so assert on the exact rendered text: ## Multi-file projects: `%load ...` and positional dirs/globs (PR #146) +### Per-file document isolation probes + +- Give only one file a root-level `private import ScalarValues::*;`. A second + file's bare `Real` must stay unresolved, as must a later prompt declaration's + bare `Real`. A qualified expression such as `%eval A::x + 1.0` should still + work, proving isolation did not remove the loaded package from the index. +- Two loaded files declaring the same root package are two root namespaces, not + a duplicate. References select the declaration in the document whose name + sorts first, independent of CLI argument order (see the CLI reference's + Multiple Files section). Put `A::X` in `first.sysml` and `A::Y` in + `second.sysml`, then reverse arguments: `A::X` must resolve and `A::Y` must + remain unresolved in both orders. Do not confuse reference precedence with + document rendering order. +- For rendering order, `%view` takes a **view**, not an ordinary package. + `%render #table` renders the loaded documents without a declared view; reverse + two nonalphabetical package names and assert their member groups reverse. + A declared view with `render asElementTable;` needs `private import Views::*;` + in its scope. +- `%save` passes notation through the formatter. Test source retention separately + from byte equality: tabs can become four spaces even while comments, members, + file order and typed declarations survive. Compare with a `develop` build + before attributing such formatting to a load-path regression. +- Both debugger fixtures in `internal/frontend/repl/testdata/` are load-ready: + `action_debug.sysml` (`%action Debug::tally`, `%step`, type `part def Z;`, + `%continue`) ends at `total = 5`; `state_debug.sysml` (`%state Debug::Cycle`, + `%advance 1`, type `part def Z;`, `%advance 9`, `%advance 5`) reaches working + at t=10 and done at t=15. This tests symbol rebinding across prompt edits. + +#### Devin Secrets Needed + +None for local multi-file CLI/REPL tests. + +### Parallel file-load verification + +- The shared concurrency knob is `-jobs N` / `OPENSYSML_JOBS`, also observable + with `%jobs`. It bounds both plan execution and files parsed/validated in one + load. Compare stdout, stderr and exit status at 1, 2 and 8 jobs plus an + environment override; test invalid values against a nonexistent path to + distinguish startup rejection from a load failure. `-workers` is not a + supported replacement flag. +- Generate a small multi-file model from the nested tools module: + `go run -C tools ./cmd/stress-model -planes 4 -satellites 10 -ground-stations 8 -split-planes `. + Verify SHA256/name manifest records, then shrink the model: unchanged surplus + output should disappear, while an edited surplus plane and user file survive. + Editing a still-current output should refuse the entire regeneration with + `nothing written`; compare all directory bytes before and after. +- A load containing `part component : Needed::T;` gives a non-vacuous batch + diagnostic. In `%verbosity debug`, declare `package Needed { part def T; }`, + then `package Needed {}`, then restore `T`. Whole-buffer diagnostics must + change 1→0→1→0. An empty package removes its members; a nonempty declaration + merges with existing members and is not a suitable deletion probe. +- For save/reload, retain comments and loaded-file declarations and assert + only the latest prompt redeclaration survives. Re-evaluate a compound + expression after `%clear` and reloading the saved file. + `sysml ...` and `%load ...` expand to model files via `internal/workspace/project.Expand`, and every file is accepted before one analysis pass -(`Session.SubmitAll`), so load order does not affect name resolution. Shapes to expect: +(`Session.SubmitAll`), each file a workspace document of its own indexed with the +others. Repeated root names resolve by document-name order, not load order. +Shapes to expect: - More than one file prints a `loaded N files:` header listing each path (a single file prints no header — a good tell that the multi-file path was taken). diff --git a/changes/unreleased/parallel-batch-validation.performance.md b/changes/unreleased/parallel-batch-validation.performance.md new file mode 100644 index 0000000000..83b41d0541 --- /dev/null +++ b/changes/unreleased/parallel-batch-validation.performance.md @@ -0,0 +1 @@ +- **A batch of files is parsed and analysed on a pool of workers.** `sysml -validate`, `-satisfy`, `-check` and `%load` open the files they are given as one batch: the files are parsed on one worker per CPU, indexed once, wildcard imports are expanded once for the batch instead of once per file, and the documents are analysed in parallel, each with a resolver and semantic model of its own over an index nothing writes meanwhile; the facts the workspace-wide audits (OOSEM, MOSA, identity metadata) need of every document are gathered once for the batch and shared by the workers. The diagnostics are the same at any job count, in command-line order; `-jobs N` or `OPENSYSML_JOBS`, the setting that already bounds how many runs of one check go concurrently, sets the pool. On an eight-CPU machine the 1 600-satellite stress constellation split over 34 files validates in 9.2 s against 22.0 s on one job and 20.3 s as a single file, within a tenth of the single file's peak memory; a 200-satellite split in 1.19 s against 2.63 s. `stress-model -split-planes ` writes the constellation one file per orbital plane, and `docs/project/satellite-network-stress-test.md` records where the split model's remaining time goes. diff --git a/changes/unreleased/per-file-documents.changed.md b/changes/unreleased/per-file-documents.changed.md new file mode 100644 index 0000000000..8ef5196eba --- /dev/null +++ b/changes/unreleased/per-file-documents.changed.md @@ -0,0 +1,2 @@ +- **Every file the command line or `%load` reads is a document of its own.** `sysml -validate`, `-satisfy`, `-e` and the REPL's `%load` used to join the files they were given into one buffer with the typed transcript, so a model split over files was analysed as if it were one file; each file is now a workspace document, indexed with the others and analysed on its own, exactly as the editor and the OMG corpus gates analyse it. Two things a reader will observe: a root-level import in one file (`private import ScalarValues::*;`) no longer serves the other files on the command line or the prompt after `%load` — a KerML root import surfaces its names in its own document's root namespace only, as `docs/project/spec-compliance.md` records — and two files that both declare `package A` are no longer reported as `Duplicate of other owned member name`: they are two root namespaces of one name, and a reference to `A` resolves to the declaration in the file whose name sorts first (the order the editor gives documents, whatever order the files were given in), as the pilot implementation resolves a repeated root name to the first. Root packages stay reachable from every file and from the prompt through the global namespace. A differential test runs every multi-file directory of the fixtures and of the four OMG corpora through the command line and through a workspace and asserts the same diagnostics. +- **A file named `` is refused by every load.** `` is the name the REPL keeps the typed transcript under, so `%load ` and `sysml -validate ` report `cannot load : the name is reserved for the text typed at the prompt` before anything loads and exit as a read failure does; `./` loads it. diff --git a/cmd/sysml/check.go b/cmd/sysml/check.go index caf9a5196b..ac2465e0f5 100644 --- a/cmd/sysml/check.go +++ b/cmd/sysml/check.go @@ -617,8 +617,8 @@ func runChecks(files []string, exprs []string, c checks) int { return rep.finish() } - // The files are loaded as one submission, indexed and analyzed once, and - // each is still summarized on its own. + // The files are loaded as one submission, each a document of its own indexed + // with the others, and each is summarized on its own. loaded, err := sess.LoadFilesSummary(paths) if err != nil { rep.failed(err.Error()) diff --git a/cmd/sysml/jobs_load_test.go b/cmd/sysml/jobs_load_test.go new file mode 100644 index 0000000000..73ac43647a --- /dev/null +++ b/cmd/sysml/jobs_load_test.go @@ -0,0 +1,83 @@ +package main + +import ( + "os" + "path/filepath" + "strings" + "testing" +) + +// A model of several files whose validation crosses them: a metadata body to +// link, a reference into another file, and an error to report in a third. +var loadModel = map[string]string{ + "a.sysml": "package A {\n\tmetadata def M { attribute n; }\n\tpart def X { @M { n = 1; } }\n}\n", + "b.sysml": "package B { part x : A::X; part y : Missing; }\n", + "c.sysml": "package C { part z : A::X { @A::M { n = 2; } } }\n", +} + +func writeLoadModel(t *testing.T) []string { + t.Helper() + dir := t.TempDir() + paths := make([]string, 0, len(loadModel)) + for _, name := range []string{"a.sysml", "b.sysml", "c.sysml"} { + path := filepath.Join(dir, name) + if err := os.WriteFile(path, []byte(loadModel[name]), 0o644); err != nil { + t.Fatal(err) + } + paths = append(paths, path) + } + return paths +} + +// TestJobsGovernLoads: -jobs and OPENSYSML_JOBS bound how many files of a load +// are parsed and validated at once; -validate reports the same at any count, and a +// bad value is rejected before any file is read, in every mode that loads a model. +func TestJobsGovernLoads(t *testing.T) { + binary := buildCLI(t) + paths := writeLoadModel(t) + validate := func(env []string, args ...string) runOutcome { + return checkPathsEnv(t, binary, env, append(append([]string{"-validate"}, args...), paths...)...) + } + want := validate(nil) + if want.status != 2 || !strings.Contains(want.output(), "unresolved reference: Missing") { + t.Fatalf("the default run should report b.sysml's unresolved name and exit 2, got %d:\n%s", want.status, want.output()) + } + + for _, jobs := range []string{"1", "2", "8"} { + if got := validate(nil, "-jobs", jobs); got.status != want.status || got.output() != want.output() { + t.Errorf("-jobs %s reported %d\n%s\nwant the default's %d\n%s", jobs, got.status, got.output(), want.status, want.output()) + } + if got := validate([]string{"OPENSYSML_JOBS=" + jobs}); got.status != want.status || got.output() != want.output() { + t.Errorf("OPENSYSML_JOBS=%s reported %d\n%s\nwant the default's %d\n%s", jobs, got.status, got.output(), want.status, want.output()) + } + } + + for _, bad := range []string{"0", "-3", "two", "1.5"} { + got := validate(nil, "-jobs", bad) + if got.status != 2 || !strings.Contains(got.output(), `-jobs="`+bad+`" is not a positive integer`) || strings.Contains(got.output(), "Missing") { + t.Errorf("-jobs %s: status %d\n%s", bad, got.status, got.output()) + } + got = validate([]string{"OPENSYSML_JOBS=" + bad}) + if got.status != 2 || !strings.Contains(got.output(), `OPENSYSML_JOBS="`+bad+`" is not a positive integer`) || strings.Contains(got.output(), "Missing") { + t.Errorf("OPENSYSML_JOBS=%s: status %d\n%s", bad, got.status, got.output()) + } + } + + if got := validate([]string{"OPENSYSML_JOBS=nope"}, "-jobs", "2"); got.status != want.status || got.output() != want.output() { + t.Errorf("-jobs 2 under OPENSYSML_JOBS=nope reported %d\n%s\nwant the default's\n%s", got.status, got.output(), want.output()) + } + + // Every mode that loads a model reads the setting before loading. + for _, mode := range [][]string{ + {"-satisfy"}, + {"-query", `sysml:name="X"`}, + {"-render", "A::X"}, + {"-render-all", t.TempDir()}, + {"-compile", "A::X", "-o", filepath.Join(t.TempDir(), "x")}, + } { + got := checkPathsEnv(t, binary, []string{"OPENSYSML_JOBS=0"}, append(mode, paths...)...) + if got.status != 2 || !strings.Contains(got.output(), `OPENSYSML_JOBS="0" is not a positive integer`) || strings.Contains(got.output(), "Missing") { + t.Errorf("%s under OPENSYSML_JOBS=0: status %d\n%s", mode[0], got.status, got.output()) + } + } +} diff --git a/cmd/sysml/load_test.go b/cmd/sysml/load_test.go index ece381b297..bacf09f2ed 100644 --- a/cmd/sysml/load_test.go +++ b/cmd/sysml/load_test.go @@ -113,11 +113,62 @@ func TestCheckGlobExitStatus(t *testing.T) { wantReport(t, checkPaths(t, binary, "-validate", filepath.Join(dir, "*.kerml")), 2, "no model files match") } +// TestLoadRefusesTheTranscriptsName checks that a file named as the session's +// transcript is a read failure like any other: refused before anything loads, +// named in the error, and exiting as a run that decided nothing. +func TestLoadRefusesTheTranscriptsName(t *testing.T) { + const reserved = "" + dir := t.TempDir() + write(t, filepath.Join(dir, reserved), "package FromFile { part def X; }\n") + write(t, filepath.Join(dir, "ok.sysml"), "package OK { part def Y; }\n") + t.Chdir(dir) + + sess := repl.NewSession() + status, err := loadFiles(sess, []string{"ok.sysml", reserved}) + var named *repl.ReservedNameError + if !errors.As(err, &named) || named.Name != reserved { + t.Fatalf("loadFiles error = %v, want a *repl.ReservedNameError naming %s", err, reserved) + } + if status != exitUnevaluable { + t.Errorf("status = %d, want %d", status, exitUnevaluable) + } + if got := sess.List(); len(got) != 0 { + t.Errorf("a refused load must load nothing, got %v", got) + } + + binary := buildCLI(t) + wantReport(t, checkPathsIn(t, dir, binary, "-validate", reserved), exitUnevaluable, + "sysml: cannot load : the name is reserved") + wantReport(t, checkPathsIn(t, dir, binary, "-constraint", "OK::Held", "ok.sysml", reserved), exitUnevaluable, + "cannot load : the name is reserved") + wantReport(t, checkPathsIn(t, dir, binary, "-validate", "./"+reserved), exitHolds) +} + // checkPaths runs the binary on paths the caller names, rather than on a model // written to a file for it as check does. func checkPaths(t *testing.T, binary string, args ...string) runOutcome { + t.Helper() + return checkPathsEnv(t, binary, nil, args...) +} + +// checkPathsEnv is checkPaths with variables added to the binary's environment. +func checkPathsEnv(t *testing.T, binary string, env []string, args ...string) runOutcome { + t.Helper() + return runPaths(t, "", binary, env, args...) +} + +// checkPathsIn is checkPaths run from dir, so a relative path is read there. +func checkPathsIn(t *testing.T, dir, binary string, args ...string) runOutcome { + t.Helper() + return runPaths(t, dir, binary, nil, args...) +} + +// runPaths runs the binary on the paths in dir with env added to its environment. +func runPaths(t *testing.T, dir, binary string, env []string, args ...string) runOutcome { t.Helper() cmd := exec.Command(binary, args...) + cmd.Dir = dir + cmd.Env = append(os.Environ(), env...) var stdout, stderr bytes.Buffer cmd.Stdout, cmd.Stderr = &stdout, &stderr err := cmd.Run() diff --git a/cmd/sysml/main.go b/cmd/sysml/main.go index d8c5495ad8..5f446db19a 100644 --- a/cmd/sysml/main.go +++ b/cmd/sysml/main.go @@ -503,6 +503,9 @@ func runCLI() int { fmt.Fprintln(os.Stderr, "sysml: -compile needs -o to name the executable (or the source file, with -source)") return 2 } + if status := resolveRunBounds(); status != 0 { + return status + } if err := runCompile(args); err != nil { return fail(err) } @@ -592,6 +595,9 @@ func runCLI() int { return refuse(modelChecks, "-render-all writes views out and decides nothing about the model; check it in its own run") } + if status := resolveRunBounds(); status != 0 { + return status + } if err := runRenderAll(args); err != nil { return fail(err) } @@ -635,6 +641,9 @@ func runCLI() int { fmt.Fprintln(os.Stderr, "sysml: -query cannot be combined with checks, -eval, -render, -render-document, -output or -from") return 2 } + if status := resolveRunBounds(); status != 0 { + return status + } return runQuery(args, queryText) } @@ -647,6 +656,9 @@ func runCLI() int { fmt.Fprintln(os.Stderr, "sysml: -render and -render-document each write a document out; ask for one per run") return 2 } + if status := resolveRunBounds(); status != 0 { + return status + } if err := runRender(args); err != nil { return fail(err) } diff --git a/cmd/sysml/usage.go b/cmd/sysml/usage.go index 123662dc2d..820a28b25e 100644 --- a/cmd/sysml/usage.go +++ b/cmd/sysml/usage.go @@ -104,7 +104,8 @@ func doc() usage.Doc { "run, so machines on sibling parts share it and its connectors.", "-jobs is how many runs of one check may go concurrently — the " + "linearizations of an exploration, the engines -engine all consults " + - "— each on a worker of its own over the shared model; the result is " + + "— each on a worker of its own over the shared model, and how many " + + "files of one load are parsed and validated at once; the result is " + "the same at any count. By default one per CPU, fewer where the memory " + "available leaves less than 512 MiB per worker.", }, @@ -510,7 +511,7 @@ func doc() usage.Doc { }, { Title: "Environment", ManOnly: true, - Items: append(append(append(usage.BudgetEnvironment(), usage.JobsEnvironment()...), usage.ToolEnvironment()...), solverEnvironment()...), + Items: append(append(append(usage.BudgetEnvironment(), usage.LoadJobsEnvironment()...), usage.ToolEnvironment()...), solverEnvironment()...), Paragraphs: []string{usage.LegacyPrefixNote, usage.BudgetScopeNote}, }, { Title: "Files", @@ -578,7 +579,7 @@ func registerFlags(fs *flag.FlagSet) { fs.Var(&modelChecks.observe, "observe", "Report this feature of the -runs action, or clock for the time it completed at; default every feature it holds and the clock. With -compare-results, the stored observable to compare, or = to read it from another feature of the run (repeatable)") fs.Var(&modelChecks.sweeps, "sweep", "Run the -analysis or -calc once per value of this range, as -sweep \"n=1..8:2\"; several ranges run their cartesian product (repeatable)") fs.Var(&modelChecks.samples, "samples", "Draw this many values uniformly from each -sweep range instead of stepping through it; needs -seed") - fs.Var(&jobsFlag, "jobs", "Runs of one check that may go concurrently, each on a worker of its own; default OPENSYSML_JOBS, else one per CPU, fewer where the memory available leaves less than 512 MiB per worker") + fs.Var(&jobsFlag, "jobs", "Runs of one check that may go concurrently, each on a worker of its own, and files of one load that are parsed and validated at once; the result is the same at any count. Default OPENSYSML_JOBS, else one per CPU, fewer where the memory available leaves less than 512 MiB per worker") fs.BoolVar(&listEngines, "engines", false, "List the analysis engines this build knows and their status, spawning nothing, and exit") fs.BoolVar(&probeEngines, "probe", false, "With -engines, also start each external engine once and report the outcome as its status") diff --git a/docs/guide/04-repl.md b/docs/guide/04-repl.md index 0b9ac00d8d..7050f43542 100644 --- a/docs/guide/04-repl.md +++ b/docs/guide/04-repl.md @@ -130,20 +130,27 @@ session, so the next submission is parsed against the model as it stood before t non-interactive use, a load's diagnostics are errors, so a script that loads a malformed file fails rather than continuing against an empty session. -Two loaded files that both open `package P` declare two packages of that name, and the load -reports this: +Each loaded file is a document of its own, analysed as the editor and the checker analyse it, +while everything typed at the prompt forms one transcript document. The transcript is kept +under the name ``, which is therefore reserved: a file whose path is literally `` +is refused by `%load` and by the command line (`cannot load : the name is reserved for the +text typed at the prompt`) before anything is loaded, like a file that could not be read; name it +`./` or from another directory to load it. Two consequences follow. + +A root-level import serves the file it is written in and no other: after `%load a.sysml`, a +`private import ScalarValues::*;` at the top of `a.sysml` does not make `Real` resolvable in +another loaded file or at the prompt. Write the import where it is used — in each file, or at +the prompt. The packages a file declares stay reachable from every other file and from the +prompt, since root packages share the global namespace. + +Two loaded files that both open `package P` declare two root packages of that name. That is not +a duplicate, and the load reports it as a note rather than a warning: ``` sysml> %load a.sysml b.sysml loaded 2 files: a.sysml b.sysml -a.sysml:1:9: warning: Duplicate of other owned member name -package P { part def A; } - ^ -b.sysml:1:9: warning: Duplicate of other owned member name -package P { part def B; } - ^ ✓ package P ✓ package P note: P is opened by more than one loaded file; each opening stays a declaration of its own, so a member of one is not visible unqualified in the other — qualify it (P::member) @@ -151,10 +158,25 @@ note: P is opened by more than one loaded file; each opening stays a declaration Each file keeps its own identity, which is what lets you reload one of them and replace only its own contribution. If the two openings were merged into a single namespace, an edit to one -file could silently delete the other file's members. The members of both openings are -declared and reachable when qualified (`P::Wheel`, `P::Axle`), but an unqualified reference -from one to the other does not resolve. Entering a package at the prompt is unaffected: it still -merges into the package already in the session. +file could silently delete the other file's members. A qualified reference to `P` resolves to +the declaration in the file whose name sorts first — the order the editor gives documents, +not the order the files were loaded in — so `P::A` resolves while `P::B` does not; an unqualified reference +from one opening to the other does not resolve either. Entering a package at the prompt is +unaffected: it still merges into the package already in the session. + +### Known behaviour + +Two evaluation rules of the prompt predate per-file loading and are kept as they are; both are +open to change. A prompt expression (`%eval`, the arguments of `%calc` and `%sweep`) evaluates in +the last namespace declared, and when a loaded file declares it, the expression sees that file's +root imports even though a typed declaration does not — after `%load a.sysml` with +`private import ScalarValues::*; package A { … }`, `%eval 1.5 as Real` resolves while a typed +`attribute y : Real;` reports `Real` unresolved; the alternative is a fallback to the transcript's +own root. A qualified command argument (`%eval A::y`, `%print A::y`) is looked up in the symbol +index, which holds every document's declarations, so with two loaded `package A` it reaches the +member of the second `A` that a reference in a model or in a compound expression cannot (`%eval A` +alone reports `A` ambiguous); the alternatives are a resolver-based lookup for the evaluating +commands, or rejecting a root that resolves ambiguously. ## Finding what a build offers diff --git a/docs/internals/performance.md b/docs/internals/performance.md index 213ff2b0af..09dd7bdb59 100644 --- a/docs/internals/performance.md +++ b/docs/internals/performance.md @@ -311,6 +311,109 @@ so the win is collector pressure rather than bytes: Diagnostics and exit status were verified byte-identical against the previous binary over the same models. +### What a batch of files costs + +`sysml -validate a.sysml b.sysml …`, `-satisfy` and `%load` open the files as +one batch of workspace documents (`model.(*Workspace).OpenAll`): the files are +parsed and their scope trees built on a pool of workers, installed in the +shared index one after another, wildcard imports are expanded once for the +batch, and the documents are analyzed on the pool +(`model.(*Workspace).DiagnosticsAll`), each in a `passes.Context` of its own +with a private resolver and semantic model, over an index nothing writes while +the pool runs. What resolving a document would otherwise link into the scope +tree on first use — the owner of a metadata body — is linked for every document +of the batch before the pool starts (`passes.PrepareBatch`), so the workers +only read it. The batch carries a `passes.Gathers` of its own +(`passes.Batch.Gathers`): the first context that runs a workspace-wide audit +gathers every document's facts into it, under its lock, and every context reads +the same union afterwards, so the audits gather each document once per batch +rather than once per analysis. A private resolver records no dependencies, so +the diagnostics a batch computes are cached with none — dropped on any change +to the workspace (`Workspace.batched`) rather than per dependency — and what +its contexts gather never enters the workspace's own `passes.Gathers`, whose +entries are invalidated per document as the editor path's resolver reports. +The batch does settle the regathers an edit left pending before it consults +the diagnostics cache, so a verdict the editor path cached is not served once a +change to another document has undone it. Diagnostics come back in the order the +files were given and are the same at any job count; `-jobs` and +`OPENSYSML_JOBS`, the setting that bounds how many runs of one check go +concurrently, set the pool, default one worker per CPU. The earlier cost of +indexing files one at a time — re-expanding wildcard imports over every +document loaded so far, quadratic in the file count — is gone with it. + +Measured on the satellite constellation split one file per orbital plane +(`stress-model -split-planes`; `Intel Xeon Platinum 8559C`, 8 CPUs, 31 GiB, +Go 1.25.0, one run each, `/usr/bin/time -v`): + +| model | files | jobs | wall | CPU | peak RSS | +| ----- | ----- | ---- | ---- | --- | -------- | +| 1 600 satellites, one file | 1 | — | 20.3 s | 133% | 2.38 GB | +| 1 600 satellites, split | 34 | 1 | 22.0 s | 129% | 2.27 GB | +| | | 2 | 14.3 s | 213% | 2.12 GB | +| | | 4 | 10.7 s | 288% | 2.23 GB | +| | | 8 | 9.24 s | 365% | 2.63 GB | +| 200 satellites, split | 10 | 1 | 2.63 s | 129% | 358 MB | +| | | 8 | 1.19 s | 341% | 507 MB | + +The split costs a little over the single file's time on one job and 2.2× less +on eight, within a tenth of the single file's peak RSS. What bounds the pool +is the gather: in the eight-job CPU profile (9.57 s wall, 34.2 s of samples) +the three audits' gather is 4.5 s — the OOSEM union 3.6 s, identity 0.52 s, +MOSA 0.37 s — run by one context over all 34 documents while the other +workers wait at the lock; the scan of the files for the root namespaces they +import from their siblings (`project.Dependencies`, 0.87 s) and installing the +scope trees and expanding wildcard imports before the pool (`commitBatch`, +1.0 s) are serial too; the rest — name resolution 9.0 s, the inherited-name +conflict pass 4.0 s, type checking 1.2 s, the collector 5.6 s — is spread over +the workers. Gathering on the pool as well, each worker gathering its own +document's facts into the union before analysis starts, is the step left to +the ~5 s the scaling design sets for this run +([scaling to very large models](../project/large-model-scaling-design.md)). +Before the audits gathered once per batch, each of the 34 analyses gathered +all 34 documents afresh in its own model: 129 s on one job, 30.9 s on eight +at 7.24 GB peak RSS, 118 s of a 128 s serial analysis in the three passes. +`BenchmarkAnalyzeSplitPerDocument` in `tests/stressmodel` measures the +pool's own speedup with the three audits left out, over the split's six files +at 512 satellites: 3.91 s → 1.09 s, one job against eight, the largest file +bounding it. The whole load of the same split, audits included, is +`BenchmarkValidateSplit`: 6.30 s → 2.14 s. + +Parallelism does not reduce what a load allocates — the 34-file run allocates +7.6 GiB and 108 million objects at any job count, against the single file's +6.1 GiB and 97 million — and the collector marking eight workers' garbage at +once is where the pool loses efficiency beyond the gather +(`runtime.gcBgMarkWorker` 16.2% and `runtime.scanobject` 17.0% of the +eight-job samples; user time 27.6 s → 32.4 s). The allocation sites the +pool does not help, from the heap profile of an earlier build's single-file 1 600-satellite +run (62 million sampled objects and 4.3 GiB, of a run that counts 85.5 million +allocations and 5.4 GiB), by objects allocated: + +| share of objects | site | what allocates | +| ---------------- | ---- | -------------- | +| 12.8% | `passes.(*w9cConflictChecker).specializes` | a slice per conformance question of the inherited-name conflict pass | +| 12.6% | `passes.contributionsOf` | the per-base member list the same pass compares | +| 9.8% | `symbols.FQNOf` (via `strings.Builder`) | a fully-qualified name built as a string, 78% of it from `symbols.(*Index).GetFQN`, 18% from the conflict pass | +| 7.3% | `resolve.(*Resolver).specializationChain` | a slice per walk of a type's generalizations | +| 3.5% | `semantics.(*Model).AllSupertypes` | a slice per supertype closure | +| 2.6% | `parser.(*Parser).parseQualifiedNameRelaxed` | a qualified-name node per reference | +| 1.9% | `parser.(*Parser).parseBase` | a node per specialization clause | + +By bytes the parser leads — `parseUsage` and what it calls are 25% of the 5.4 +GiB, `parseQualifiedNameRelaxed` alone 5% — with `contributionsOf` (6.6%), +`specializes` (4.9%) and `FQNOf` (4.4%) behind it. Each of these is one +allocation per token, per name or per lookup where one per file, or none, would +serve — the snapshot decoder's node table, allocated as one block, is the +model — and each is to be measured on its own before it is changed. + +One parse per load is spent twice: the REPL parses each file to accept it (the +names it declares, whether it closes its own text) and the workspace parses the +same bytes again as the document. The 34 files of the split (17 MB) parse in +1.8 s serially (`repl.preparse` in the one-job profile), so the second parse +is ~1.8 s of the one-job 22.0 s and ~0.3 s of the eight-job wall, where it runs +on the pool. Carrying the accepted tree into the workspace +batch would recover it; it is a change to what `model.Input` owns and is left +to be measured on its own. + ## What a process pays before the model Every `sysml`, `sysml-lsp` and `sysml-grpc` start, and every test that builds a @@ -456,12 +559,11 @@ constellation, one file, three runs each: At 1 600 satellites: 17.7 s and 5.5 GiB allocated became 20.5 s and 5.8 GiB. The cost falls on whatever analyzes through the workspace's own context: the -LSP server, a REPL session, and `sysml -validate`, which loads through a REPL -session. A batch that analyzes each document in a private `passes.Context` — -its own resolver and model over the read-only index, as a pool of workers -must — has no frames to record into and pays none of it; the workspace's -gathered facts are what such a batch should hand its workers, so that they do -not gather per worker what the workspace gathered once. +LSP server, and a REPL session's typed submissions. `sysml -validate` and +`%load` analyze each file in a private `passes.Context` — its own resolver and +model over the read-only index, as a pool of workers must — with the audits' +facts gathered once for the batch ("What a batch of files costs"), and pay +none of it. ## Notes for further work @@ -470,17 +572,6 @@ not gather per worker what the workspace gathered once. immutable once its index is built, so whether it declares any `about` usages — and which — is computable once at library-index build time; a session would then walk only workspace documents. -- Validating many files in one `sysml -validate` invocation is quadratic in - the file count: the CLI submits files one at a time and every submission - reindexes the workspace, re-running wildcard-import expansion over every - document loaded so far. The 100-file OMG training corpus costs 6.7 s as one - batch where its two halves cost 0.6 s and 3.4 s separately, and a CPU - profile of the batch spends 52% under `model.(*Workspace).setOpenBuffer` → - `reindexLocked` with `symbols.(*Index).ExpandWildcardImports` the largest - component. Submitting a batch as one indexing unit, or expanding wildcard - imports incrementally for documents a new submission cannot affect, would - make a batch cost what its parts cost. - - Runs over a large model spend their time in collection, not in the executor (above). Reducing what a load leaves behind is the lever, since the live model is what each cycle scans. diff --git a/docs/project/satellite-network-stress-test.md b/docs/project/satellite-network-stress-test.md index b3084075ef..3ee284c264 100644 --- a/docs/project/satellite-network-stress-test.md +++ b/docs/project/satellite-network-stress-test.md @@ -134,51 +134,97 @@ report for the synthetic model there, because most of their elements are attribute redefinitions with a literal value rather than definitions with bodies of their own. -## Running: instantiation, state machines and satisfaction - -`sysml -satisfy -memstats` loads and validates the model, then for every -`satisfy` assertion instantiates its subject — a satellite's configured usage, -with its subsystem and component tree — starts the mode machine the spacecraft -exhibits, evaluates the summed mass and power over the instance and checks the -requirement's constraint against it. Three assertions per satellite; every one -holds. - -| satellites | assertions | wall | of which load | allocated | peak RSS | -| ---------- | ---------- | ---- | ------------- | --------- | -------- | -| 2 | 6 | 0.10–0.15 s | 0.06 s | 65 MiB | 90 MB | -| 10 | 30 | 0.23–0.25 s | 0.13 s | 122 MiB | 118 MB | -| 50 | 150 | 0.90–0.98 s | 0.49 s | 402 MiB | 210 MB | -| 100 | 300 | 1.86–1.95 s | 0.95 s | 762 MiB | 310 MB | -| 200 | 600 | 3.9–4.1 s | 2.0 s | 1.5 GiB | 530 MB | -| 400 | 1 200 | 8.3 s | 4.5 s | 2.9 GiB | 975 MB | -| 800 | 2 400 | 17.5–17.9 s | 8.7 s | 6.0 GiB | 1.83 GB | -| 1 600 | 4 800 | 38.1 s | 19.0 s | 12.8 GiB | 3.8 GB | -| 3 200 | 9 600 | 83 s | 42 s | 28.9 GiB | 7.6 GB | - -Checking the whole constellation costs **about 2.0× a validation** of the -same model at every size, and the extra is linear: about 3.2 ms, 1.2 MiB -allocated and 0.45 MB of peak RSS per assertion — that is, per instantiation -of a satellite with its twenty components and a running state machine. A CPU -profile at 400 satellites puts the run's own share (30% of samples, the rest -being the load) almost entirely in `runtime.(*Context).Instantiate`: -materializing the parts that run behaviors (`materializeBehavingParts`, -`runsBehaviors`) and shaping the features of each type (`FeaturesOf`, -`semantics.(*Model).ShapeFeatures`). Evaluating the budgets is a small part. - -Re-checking a loaded constellation is much cheaper than the first check, -because the runtime's per-type memoization — feature shapes, which types run -behaviors, the verification cases under each scope — is then warm. -`BenchmarkSatisfy` measures the warm re-check of every assertion: - -| satellites | assertions | warm re-check | allocated | -| ---------- | ---------- | ------------- | --------- | -| 32 | 96 | 3.5 ms | 1.9 MiB | -| 128 | 384 | 20 ms | 11.2 MiB | -| 512 | 1 536 | 144 ms | 104 MiB | - -That is under 0.1 ms per assertion warm, against 3.2 ms cold: the first check -pays for building the runtime's view of every type, and a session that keeps -the model loaded — the REPL, the gRPC service — amortizes it. +### Split by plane, parallel + +A project of this size is not one file. `stress-model -split-planes ` +writes the same constellation as one `.sysml` per orbital plane plus +`library.sysml` (the definitions every plane shares) and `constellation.sysml` +(the ground segment and the cross-plane network); the split declares the same +network and analyzes to the same diagnostics as the single file +(`TestSatelliteNetworkSplitValidates`). The generator lists what it wrote in +`.stress-model-files` beside the model, each with a digest of its content, and +a later generation into the same directory replaces only files on that list +that still read as written, removes those of them it did not write again, and +writes nothing at all when a file it would write is not on the list, has been +edited or is not a regular file — so a smaller constellation leaves no plane of +a larger one behind and nothing else in the directory, a file of the user's or +an edited plane, is touched. `sysml -validate` over the files parses +them on a pool of workers, indexes them once, expands wildcard imports once +and analyzes them on the pool, each document with a resolver and semantic +model of its own; the facts the workspace-wide audits (OOSEM, MOSA, identity +metadata) need of every document are gathered once for the batch and read by +every worker. `-jobs N` (or `OPENSYSML_JOBS`), the same setting that bounds +how many runs of one check go concurrently, sets the pool, default one worker +per CPU. The diagnostics are the same at any job count, in command-line order. + +```bash +go run -C tools ./cmd/stress-model -planes 32 -satellites 50 -ground-stations 160 -split-planes constellation/ +/usr/bin/time -v sysml -validate -memstats -jobs 8 constellation/*.sysml +``` + +Same machine as above (`Intel Xeon Platinum 8559C`, 8 CPUs, 31 GiB, Go +1.25.0, Linux); one run per row; *CPU* is `(user + system) / wall`. + +| model | files | jobs | wall | user | CPU | allocated | peak RSS | +| ----- | ----- | ---- | ---- | ---- | --- | --------- | -------- | +| 200 satellites, one file | 1 | — | 2.15 s | 2.6 s | 126% | 806 MiB | 393 MB | +| 200 satellites, split | 10 | 1 | 2.63 s | 3.3 s | 129% | 957 MiB | 358 MB | +| | | 2 | 1.69 s | 3.3 s | 201% | 972 MiB | 379 MB | +| | | 4 | 1.29 s | 3.3 s | 271% | 974 MiB | 414 MB | +| | | 8 | 1.19 s | 3.9 s | 341% | 976 MiB | 507 MB | +| 1 600 satellites, one file | 1 | — | 20.3 s | 26.1 s | 133% | 6.1 GiB | 2.38 GB | +| 1 600 satellites, split | 34 | 1 | 22.0 s | 27.6 s | 129% | 7.5 GiB | 2.27 GB | +| | | 2 | 14.3 s | 29.5 s | 213% | 7.6 GiB | 2.12 GB | +| | | 4 | 10.7 s | 29.9 s | 288% | 7.6 GiB | 2.23 GB | +| | | 8 | 9.24 s | 32.4 s | 365% | 7.6 GiB | 2.63 GB | + +Three things the table says: + +- **Splitting the file costs little, and the pool pays it back.** One file + validates in 20.3 s; the same model in 34 files takes 22.0 s on one job + — the split adds the per-plane packages, their imports, the sibling-file + dependency scan and the gather of 34 documents instead of one — and 9.24 s + on eight, 2.2× the single file's speed. The 200-satellite split goes from + 2.63 s to 1.19 s. Peak RSS stays within a tenth of the single file's: + 2.63 GB at eight jobs against 2.38 GB, since the workers share one gather + and hold only their own document's memoization. +- **The gather is what bounds the pool.** Eight jobs reach 365% CPU, not + 700%. A CPU profile of the eight-job run (9.57 s wall, 34.2 s of samples) + puts 4.5 s in the three audits' gather — the OOSEM union 3.6 s, identity + 0.52 s, MOSA 0.37 s — which the first context to ask runs over all 34 + documents while the other workers wait for it; before the pool, the scan + of the files for the root namespaces they import from siblings + (`project.Dependencies`, 0.87 s) and installing the 34 scope trees in the + index and expanding wildcard imports (`commitBatch`, 1.0 s) are serial + too. Over 6 s of the 9.24 s is therefore on one thread. Gathering the + documents on the pool as well — each worker gathering its own document's + facts into the union, in a context of its own, before analysis starts — + would take most of the 4.5 s off the critical path and is the remaining + step to the ~5 s the [scaling design](large-model-scaling-design.md) sets + for this run. +- **The analysis itself parallelizes.** Outside the gather the profile is + the per-document work: name resolution 9.0 s of the 34.2 s, the + inherited-name conflict pass 4.0 s, type checking 1.2 s, the collector + marking eight workers' allocations at once 5.6 s. On one job the 34 + analyses are 17.1 s of samples, 3.9 s of them the gather, so a document + averages 0.39 s: no one file is a straggler that would bound the pool the + way the gather does. `BenchmarkAnalyzeSplitPerDocument` in + `tests/stressmodel` analyzes the split's six files over one index with + the three audits left out, the pool's own speedup: 3.91 s → 1.09 s at 512 + satellites on one job versus eight, the largest file about a quarter of + the work. `BenchmarkValidateSplit`, the whole load audits included, runs + 6.30 s → 2.14 s. + +**What the gather cost before.** The three audits used to gather every +workspace document once per document analyzed, each analysis in a model of +its own, so the gather was done 34 times over 34 documents: the 34-file split +took 129 s on one job and 30.9 s on eight (658% CPU, 39.1 GiB allocated, +7.24 GB peak RSS for eight workspace-wide memoizations held at once), against +18.5 s for the single file of that build. The profile of that run spent 75% of +its samples in `OOSEMMethodPass`, `IdentityMetadataPass` and `MOSAPass`; the +per-document work was about 10 s of 128 s on one job. The per-document gather +cache (`passes.Gathers`) removed the quadratic term for the editor path, and +handing one such gather to a batch's workers removed it for the command line. ## Editing: what an editor pays per keystroke @@ -324,11 +370,9 @@ the interactive band at every operation measured. usage cost more per element, long documentation comments cost less — but the shape of the curve (linear load, memory-bound batch, workspace-bound editing) does not depend on the regularity. -- The whole constellation is one file except where a figure says it was - split. Splitting it over files changes two things: the CLI submits files one - at a time and reindexes after each, which is quadratic in the file count - (`docs/internals/performance.md`, notes for further work), and an editor pays - the per-file analysis once per open file. +- Most figures are for the whole constellation as one file. The split by + plane is measured above at two sizes only; an editor also pays the per-file + analysis once per open file. - `-satisfy` instantiates each satellite's tree on its own; it does not instantiate the whole `Network` as one object with 12 800 satellites and their links, and no figure here says what that would cost. @@ -346,18 +390,13 @@ validation, and one definition with many occurrences — is in [scaling to very large models](large-model-scaling-design.md). The three items below are the ones the profiles point at directly. -- **Hand the gathered facts to the batch pipeline.** The OOSEM, MOSA and - identity audits now gather each workspace document once and judge each - analyzed document over the union, which is what made the 34-file split load - in 18 s rather than 126 s through one workspace. A batch that analyzes - documents on parallel workers with private contexts gathers per worker - again unless the workspace's gathers are what the batch hands them. - **Make a one-shot validation skip the bookkeeping.** The persistent model records, on every memoized read, which document depends on the entry's owner, so that the owner's replacement invalidates the reader. A validation that will never edit pays that for nothing — about a sixth of its wall time - at 200 satellites. Analyzing each document in a private context over the - read-only index, as a parallel batch does, records nothing. + at 200 satellites. `sysml -validate` already analyzes each document in a + private context over the read-only index and records nothing; a consumer + validating once through `Workspace.Diagnostics` still pays it. - **Reduce allocation per element.** Nineteen KiB allocated per element against 2.7 KiB held means a load produces seven times its own weight in garbage, and the collector's quarter of the profile is the price. The diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 1ec25a04e5..bb4ef2db3d 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -89,6 +89,18 @@ Load multiple files before evaluating: sysml -e "result" types.sysml instances.sysml ``` +Every file named on the command line is a document of its own, analysed as the editor and the +corpus gates analyse it, and the files are indexed together so that one file's reference to a +package another declares resolves. Two consequences follow: + +- A root-level import serves only the file it is written in. `private import ScalarValues::*;` + at the top of `types.sysml` does not make `Real` resolvable in `instances.sysml`; each file + imports what it uses. +- Two files that both declare `package A` are two root packages of that name, not a duplicate. + A reference to `A` resolves to the declaration in the file whose name sorts first (the + order the editor and the workspace give documents, whatever order the files were given + in), so `A::x` resolves where `x` is a member of that declaration. + ## Real-World Examples ### 1. Quick Calculation @@ -289,7 +301,7 @@ written in, so the verdicts are about that object: | `-engines` | Lists the analysis engines this build knows — name, kind, protocol, authority, the question kinds each answers and its status — and exits, without a model and without starting a process: the external engines of `OPENSYSML_ENGINES` and the tools of `OPENSYSML_TOOLS` are listed from their manifests alone, each followed by a line naming its file and command. See [Analysis engines](#analysis-engines) | | `-probe` | With `-engines`, also start each external engine once, check its `describe` against its manifest entry field by field and report the outcome as its status (`ready (…; describe agrees)`, or the first field that disagrees). See [External engines](external-engines.md) | | `-engine \|auto\|all` | The analysis engine every check of the invocation is put to. `auto` (the default) picks the engine of highest authority covering the question and advances past one that refuses or answers *not covered*, reaching an external engine only after every built-in one has; a name (`run`, `explore`, `check`, `smt`, `sweep`, `solve`, or an external engine's) puts the question to that engine alone, and its refusal is the answer; `all` puts it to every engine covering it, one after another in name order, and composes their answers. A name no engine is registered under is refused before anything runs. `-engine explore` explores as `-schedule explore` does; `-engine check` searches every schedule of each `-action` for a violation, a deadlock, a failure or a divergence ([Checking every schedule of an action](#checking-every-schedule-of-an-action-or-a-state-machine)); `-engine smt` decides a `-check-property` over every schedule and every value of the free inputs with an SMT solver ([Deciding a property over the inputs](#deciding-a-property-over-the-inputs)). See [Analysis engines](#analysis-engines) | -| `-jobs ` | Runs of one check that may go concurrently — the linearizations of an exploration, the rows of a `-sweep`/`-samples`, the engines `-engine all` consults — each on a worker of its own over the shared model. `n` is a positive integer; the default is `OPENSYSML_JOBS`, else one per CPU, fewer where the memory available leaves less than 512 MiB per worker (Linux: `MemAvailable` and the cgroup's `memory.max`; one at least). The result of a check is the same at any count: the outcome table, the witness, the run count and the cut a violation makes are those of the runs taken one at a time in plan order. See [Running in parallel](#running-in-parallel) | +| `-jobs ` | Runs of one check that may go concurrently — the linearizations of an exploration, the rows of a `-sweep`/`-samples`, the engines `-engine all` consults — each on a worker of its own over the shared model, and how many files of one load are parsed and validated at once (`-validate`, `-satisfy`, `-check` and every mode that loads files). `n` is a positive integer; the default is `OPENSYSML_JOBS`, else one per CPU, fewer where the memory available leaves less than 512 MiB per worker (Linux: `MemAvailable` and the cgroup's `memory.max`; one at least). The result of a check is the same at any count: the outcome table, the witness, the run count and the cut a violation makes are those of the runs taken one at a time in plan order. See [Running in parallel](#running-in-parallel) | | `-json` | Reports the checks as one JSON document rather than as lines. Each check carries its `plan` and `results[]` beside the fields it always carried ([Analysis engines](#analysis-engines)) | Other modes, each described in full by `sysml -help` and the manual page: @@ -1506,6 +1518,12 @@ the rows, their outputs, verdicts, evaluations and errors are those of `-jobs 1` With `-json` the check's `plan` carries `workers`, how many workers the plan built, and `warming`, the milliseconds spent building them; the human-readable report does not print them. +The same count sets how many files of one load are parsed and validated at once. The files named +on the command line are parsed on `n` workers, indexed together once, and analysed on `n` workers, +each file as a document of its own; the diagnostics are those of `-jobs 1`, in command-line order, +whatever `n` is. A model split over several files therefore validates faster on more CPUs where a +single file does not; `docs/project/satellite-network-stress-test.md` records the measurements. + ## Analysis engines Every check is a question put to an analysis engine, and every verdict line is followed by its diff --git a/docs/reference/environment.md b/docs/reference/environment.md index 2147929818..47e0067940 100644 --- a/docs/reference/environment.md +++ b/docs/reference/environment.md @@ -13,7 +13,7 @@ run that would never finish into a reported error instead of a hang. | `OPENSYSML_MAX_ELEMENTS` | `1000000` | Collection elements one evaluation may hold — the bound on the memory a run holds rather than on the work it does | | `OPENSYSML_MAX_CALC_DEPTH` | `10000` (ceiling `25000`) | Nested `calc` invocations one run may hold on the stack, which is what a recursion spends | | `OPENSYSML_MAX_SWEEP_RUNS` | `1000` | Runs one parameter sweep or sample may make (`-sweep`/`-samples`, `%sweep`/`%samples`, `RunSweep`), each a whole analysis or calc run with the budgets above of its own | -| `OPENSYSML_JOBS` | one per CPU, fewer where the memory available leaves less than 512 MiB per worker | Runs of one check that may go concurrently (`-jobs`, `%jobs`; the gRPC service reads it at startup), each on a worker of its own over the shared model. Bounds how many runs go at once, not the work or memory of any one of them: a fleet of `n` workers may hold `n` times `OPENSYSML_MAX_ELEMENTS`. The result of a check does not depend on it | +| `OPENSYSML_JOBS` | one per CPU, fewer where the memory available leaves less than 512 MiB per worker | Runs of one check that may go concurrently (`-jobs`, `%jobs`; the gRPC service reads it at startup), each on a worker of its own over the shared model, and how many files of one load are parsed and validated at once. Bounds how many runs go at once, not the work or memory of any one of them: a fleet of `n` workers may hold `n` times `OPENSYSML_MAX_ELEMENTS`. The result of a check does not depend on it | | `OPENSYSML_CALC_COMPILE` | unset (on) | Set to `0`, `false`, `off` or `no` to run every `calc` on the reference evaluator, instead of compiling a pure scalar body to a closure fast path on its first invocation; results, errors and step counts are the same either way, so this is a bisecting aid | | `OPENSYSML_SMT` | unset (look for `z3`, then `cvc5`, on `PATH`) | Executable the `smt` and `solve` engines (`-engine smt`, `%engine smt`) and `%check`, `%explain`, `%solve`, `%configure` and `%optimize` drive as their SMT solver, speaking SMT-LIB2 on standard input (experimental); `%optimize` needs `z3` in particular, as `(minimize …)`/`(maximize …)` is a z3 extension cvc5 does not implement | | `OPENSYSML_SMT_TIMEOUT` | `10s` | How long one solver query may take, as a Go duration (`5s`, `500ms`), after which the verdict is `unknown`; a check's `-check-timeout` (`%check-bounds timeout=`) takes its place for the `smt` engine's queries | @@ -187,7 +187,7 @@ ranges would make more runs than it allows is refused before the first one is made, naming the count the plan asks for and the bound it exceeds. `OPENSYSML_JOBS` is no budget at all but the width of the fleet: how many of one check's runs — an exploration's linearizations, a sweep's rows, the engines `-engine all` consults — -may go at once. A value that is not a positive integer is refused at startup; `-jobs` and `%jobs` +may go at once, and how many files of one load are parsed and validated at once. A value that is not a positive integer is refused at startup; `-jobs` and `%jobs` override it for one invocation or session. See [Running in parallel](cli.md#running-in-parallel). diff --git a/docs/reference/repl-commands.md b/docs/reference/repl-commands.md index 35fcf2cd50..fe566436d6 100644 --- a/docs/reference/repl-commands.md +++ b/docs/reference/repl-commands.md @@ -43,7 +43,7 @@ into the parts it holds (`car.fl.hub`, `#3.fl`, `car.wheels[2]`). | `%schedule []` | Show or set the scheduling policy the executors resolve their [choice points](../guide/06-behavior.md) under: `reverse` (the default: reverse token order, first holding guard, first enabled transition), `declared` (spawn and declaration order), `seed:` (a pseudo-random order the non-negative integer `n` fixes, so the same seed replays the same run) or `replay:` (the choice lines of a witness, followed move for move and then `reverse`'s picks one token a step — a header of `no choice points` follows the one run there is; a move the run cannot make is a `replay refused` error naming it — [Running one witness again](../guide/06-behavior.md#running-one-witness-again)). Applies to runs started from then on — `%action`, `%state`, `%analysis`; a calc's body performs nothing, so `%calc` has no choice to make — while a debugging session already under way keeps the policy it started with; every choice point a run reaches is reported and the `took …` of each `choice` line is what the policy took. A spelling naming no policy (an unknown name, `seed` or `seed:` without a number, `seed:-1`, `seed:abc`, a malformed `explore:` option, `replay:` without a readable file of choice lines, one that is empty or has a line spelling no choice) is refused and the policy is left as it was. `explore[:runs=N,depth=D]` is refused at the prompt too, as a typed error saying why: it replays a behavior from the start once per linearization, which `%action` and `%state`, stepping one run, cannot do — run `sysml -schedule explore -action ` (or `-state`, `-analysis`, `-calc`) for the outcome table ([Exploring every linearization](cli.md#exploring-every-linearization)), or send a request with that `schedule` over the wire | | `%strict [on\|off]` | Show or set strict conformance: report notation no SysML v2 production admits as an error, and reprint the session's diagnostics under the new mode ([Strict conformance](../guide/03-command-line.md#strict-conformance)) | | `%budget` | Show the five bounds one run may spend, each with the variable that raises it | -| `%jobs []` | Show or set how many runs of one check asked from then on may go concurrently — the linearizations a check explores under `%schedule explore`, the engines `%engine all` consults — the rows of a `%sweep` or `%samples` — each on a worker of its own over the session's model; `OPENSYSML_JOBS`, else one per CPU the memory available allows, until set. The result of a check is the same at any count. The count bounds a plan's runs, not the session: the held context, its objects and a debugging session under way are untouched by setting it. A value that is not a positive integer is refused and the count left as it was ([Running in parallel](cli.md#running-in-parallel)) | +| `%jobs []` | Show or set how many runs of one check asked from then on may go concurrently — the linearizations a check explores under `%schedule explore`, the engines `%engine all` consults — the rows of a `%sweep` or `%samples` — each on a worker of its own over the session's model — and how many files of one `%load` are parsed and validated at once; `OPENSYSML_JOBS`, else one per CPU the memory available allows, until set. The result of a check is the same at any count. The count bounds a plan's runs, not the session: the held context, its objects and a debugging session under way are untouched by setting it. A value that is not a positive integer is refused and the count left as it was ([Running in parallel](cli.md#running-in-parallel)) | | `%engines [probe]` | List the analysis engines of the build in name order — the kind of each, the protocol it is spoken by, the authority it carries, the question kinds it answers and its status (`ready`, `ready (z3 at …)` for one whose process was found, `unavailable: `), then one line per manifest entry naming its file and command — as the CLI's [`-engines`](cli.md#analysis-engines) does, starting nothing. `%engines probe` also starts each [external engine](external-engines.md) once, checks its `describe` against its manifest entry field by field and reports the outcome as its status, as `-engines -probe` does; any other argument is refused | | `%engine [\|auto\|all]` | Show or set the analysis engine every question asked from then on — `%constraint`, `%requirement`, `%satisfy`, `%validate`, `%calc`, `%analysis`, `%sweep`, `%samples`, `%check` and the other solver commands — is put to. `auto` (the default) picks the engine of highest authority covering the question and advances past one that refuses or answers *not covered*; a name puts it to that engine alone, whose refusal is then the verdict; `all` puts it to every covering engine, one after another in name order, and composes their answers, naming a disagreement in the interpreter's favor. Every verdict is followed by a `standing:` line — the claim, the strength of the evidence (*not covered*, *observed*, *witnessed*, *bounded*, *proved*) and what earned it — and under `all` each engine's part. A name no engine is registered under is refused and the selection left as it was. `explore` is refused at the prompt as `%schedule explore` is, since the debuggers step one run; the `%action` and `%state` debuggers keep the schedule `%schedule` set whatever the engine ([Analysis engines](cli.md#analysis-engines)). `%engine check` is the one selection that changes what `%action` does by itself: it puts the action to the `check` engine, which searches every schedule for a violation, a deadlock, a failure or a divergence and prints the verdict, instead of starting a debugging session; `%engine all` does the same, the exploration beside the checker, once a `%check-*` setting is made ([Checking every schedule](#checking-every-schedule-of-an-action-or-a-state-machine)) | | `%check-property [...\|off]` | Show or set the constraints and requirements the `check` engine evaluates at every stable state of a checked action, on its performing object where there is one; `off` (the default) names none | diff --git a/internal/check/edit/edit.go b/internal/check/edit/edit.go index 1ad3c34ed7..f56512526f 100644 --- a/internal/check/edit/edit.go +++ b/internal/check/edit/edit.go @@ -258,6 +258,7 @@ func (r *reindexer) analyzedIn(sf *source.SourceFile, root *ast.RootNamespace) * if r.indexed != nil { r.indexed(r.idx, sf, root) } + r.idx.ExpandWildcardImports() return r.idx } diff --git a/internal/check/passes/analyze.go b/internal/check/passes/analyze.go index 8cdcda5d41..cd6cf6ba30 100644 --- a/internal/check/passes/analyze.go +++ b/internal/check/passes/analyze.go @@ -7,6 +7,7 @@ import ( "github.com/Open-MBEE/OpenSysML/internal/check/passes/diagram" "github.com/Open-MBEE/OpenSysML/internal/check/passes/document" "github.com/Open-MBEE/OpenSysML/internal/check/passes/identity" + "github.com/Open-MBEE/OpenSysML/internal/semantic/resolve" "github.com/Open-MBEE/OpenSysML/internal/semantic/symbols" "github.com/Open-MBEE/OpenSysML/internal/syntax/ast" "github.com/Open-MBEE/OpenSysML/internal/syntax/diag" @@ -122,6 +123,32 @@ func AnalyzeWithOptions(name string, kind source.Kind, root *ast.RootNamespace, return analyze(NewContextWithOptions(name, kind, idx, parseDiags, opts), root) } +// PrepareBatch links what resolving every workspace document would write into +// its scope tree, so AnalyzeInBatch contexts only read the index. Call it alone, +// first. Every document, not only the batch's: the workspace-wide gathers read +// them all. The linker resolves as a context does, model attached. +func PrepareBatch(idx *symbols.Index, batch *Batch) { + if idx == nil || batch == nil { + return + } + linker := resolve.New(idx) + model := NewTypedModel(linker) + linker.SetModel(model) + model.SetSourceText(batch.Source) + for _, name := range idx.WorkspaceDocuments() { + linker.LinkMetadataBodies(name) + } +} + +// AnalyzeInBatch validates one document of a prepared batch in a context of its +// own; the result does not depend on which documents share the batch. +func AnalyzeInBatch(name string, kind source.Kind, root *ast.RootNamespace, + parseDiags []diag.Diagnostic, idx *symbols.Index, opts Options, batch *Batch) []diag.Diagnostic { + ctx := NewContextWithOptions(name, kind, idx, parseDiags, opts) + ctx.InBatch(batch) + return analyze(ctx, root) +} + // AnalyzeShared validates a document over a resolver and model kept across // analyses: what the run memoizes is owned by the document (see // Resolver.InDocument), to be dropped when it or what it read changes. diff --git a/internal/check/passes/batch_test.go b/internal/check/passes/batch_test.go new file mode 100644 index 0000000000..d9c8e23f47 --- /dev/null +++ b/internal/check/passes/batch_test.go @@ -0,0 +1,99 @@ +package passes + +import ( + "reflect" + "slices" + "testing" + + "github.com/Open-MBEE/OpenSysML/internal/semantic/symbols" + "github.com/Open-MBEE/OpenSysML/internal/syntax/ast" + "github.com/Open-MBEE/OpenSysML/internal/syntax/parser" + "github.com/Open-MBEE/OpenSysML/internal/syntax/source" +) + +// preparedBatch is a batch whose annotation types are reached every way member +// lookup can go: declared, inherited, through an inherited import, a filtered +// import, a typed usage, and across documents. +var preparedBatch = map[string]string{ + "meta.sysml": `package Meta { + metadata def Tag { attribute n; } + metadata def Other { attribute m; } +}`, + "model.sysml": `package P { + part def Base { public import Meta::*; } + part def Sub :> Base; + part def Filtered { public import Meta::*[@Meta::Tag]; } + part b : Base; + part def Owned { metadata def Own { attribute n; } } + part def Derived :> Owned; + part def C { + @Meta::Tag { n = 1; } + @Sub::Tag { n = 2; } + @Filtered::Tag { n = 3; } + @b::Tag { n = 4; } + @Derived::Own { n = 5; } + @Missing { n = 6; } + } +}`, +} + +func indexedBatch(t *testing.T, docs map[string]string) (*symbols.Index, map[string]*ast.RootNamespace) { + t.Helper() + idx := symbols.NewIndex() + roots := make(map[string]*ast.RootNamespace, len(docs)) + for name, src := range docs { + p := parser.New(source.New(name, []byte(src))) + root := p.ParseFile() + if len(p.Diagnostics) != 0 { + t.Fatalf("%s: parse diagnostics %v", name, p.Diagnostics) + } + idx.AddDocument(name, root) + roots[name] = root + } + idx.ExpandWildcardImports() + return idx, roots +} + +// ownersOf renders the owner of every scope of the tree in tree order, "" for none. +func ownersOf(idx *symbols.Index, root *symbols.Scope) []string { + var out []string + var visit func(*symbols.Scope) + visit = func(s *symbols.Scope) { + owner := "" + if s.Owner() != nil { + owner = idx.GetFQN(s.Owner()) + } + out = append(out, owner) + for _, c := range s.Children() { + visit(c) + } + } + visit(root) + return out +} + +// Preparing a batch links exactly the annotation bodies analyzing its documents +// links, wherever the annotation type is found, so analysis writes nothing. +func TestPrepareBatchLinksWhatAnalysisLinks(t *testing.T) { + const name = "model.sysml" + analyzed, roots := indexedBatch(t, preparedBatch) + AnalyzeWithOptions(name, source.KindSysML, roots[name], nil, analyzed, Options{}) + want := ownersOf(analyzed, analyzed.DocumentRoot(name)) + if !slices.ContainsFunc(want, func(o string) bool { return o != "" }) { + t.Fatal("analysis linked no annotation body, so the comparison proves nothing") + } + + prepared, _ := indexedBatch(t, preparedBatch) + PrepareBatch(prepared, &Batch{Documents: []string{"meta.sysml", name}}) + if got := ownersOf(prepared, prepared.DocumentRoot(name)); !reflect.DeepEqual(got, want) { + t.Errorf("preparing links owners\n%q\nwant those analysis links\n%q", got, want) + } + + // The gathers read every workspace document, so a batch of one document + // prepares the others' bodies too. + others, _ := indexedBatch(t, preparedBatch) + PrepareBatch(others, &Batch{Documents: []string{"meta.sysml"}}) + if got := ownersOf(others, others.DocumentRoot(name)); !reflect.DeepEqual(got, want) { + t.Errorf("preparing a batch without %s links its owners\n%q\nwant those analysis links\n%q", name, got, want) + } +} diff --git a/internal/check/passes/kit/pass.go b/internal/check/passes/kit/pass.go index 8bfd35a09e..5412d68b2e 100644 --- a/internal/check/passes/kit/pass.go +++ b/internal/check/passes/kit/pass.go @@ -55,6 +55,9 @@ type Context struct { // Options is what the caller asked for, fixed at construction: a pass reads // it, and nothing mutates it during a run. Options Options + // Batch is the batch this document is analyzed in, nil when it is analyzed + // alone; every context of a batch reads the one value and none writes it. + Batch *Batch resolver *resolve.Resolver model *semantics.Model @@ -67,6 +70,19 @@ type Context struct { failures []source.Span } +// Batch is what a batch of analyses computes once before its documents are +// analyzed together; anything a pass would gather over every document belongs here. +type Batch struct { + // Documents names the documents the batch analyzes, in the order asked for. + Documents []string + // Gathers is what the workspace-wide audits gather, once for the batch, on + // first use by any of its contexts; nil leaves each context to gather alone. + Gathers *Gathers + // Source reads the documents' notation, which comment and documentation + // bodies come from; nil leaves every body unreadable, as an editor never is. + Source source.Lookup +} + // Options is the analysis configuration of one run. The zero value is what // every existing caller gets: today's behavior, unchanged. type Options struct { @@ -85,6 +101,15 @@ func (c *Context) Share(resolver *resolve.Resolver, model *semantics.Model, gath c.resolver, c.model, c.gathers = resolver, model, gathers } +// InBatch places the context in batch, reading the gathers the batch shares +// instead of gathering alone; a nil batch leaves it analyzing on its own. +func (c *Context) InBatch(b *Batch) { + c.Batch = b + if b != nil { + c.gathers = b.Gathers + } +} + // Gathers is what the workspace-wide audits gathered per document, shared // across analyses by a workspace and made afresh for a context outside one. func (c *Context) Gathers() *Gathers { @@ -141,6 +166,9 @@ func (c *Context) Model() *semantics.Model { c.model = c.newModel(c.Resolver()) // Attach model to resolver for inheritance-aware member resolution c.Resolver().SetModel(c.model) + if c.Batch != nil { + c.model.SetSourceText(c.Batch.Source) + } } return c.model } diff --git a/internal/check/passes/pass.go b/internal/check/passes/pass.go index cfd6ac540f..7dcb1283c2 100644 --- a/internal/check/passes/pass.go +++ b/internal/check/passes/pass.go @@ -28,6 +28,9 @@ type Context = kit.Context // Options is the analysis configuration of one run. type Options = kit.Options +// Batch is what a batch of analyses shares: its documents, gathers and source lookup. +type Batch = kit.Batch + // NewContext builds a Context for a document, in the default mode. func NewContext(name string, idx *symbols.Index, parseDiags []diag.Diagnostic) *Context { return NewContextWithKind(name, source.KindOf(name), idx, parseDiags) diff --git a/internal/frontend/repl/analysis.go b/internal/frontend/repl/analysis.go index 39e8483fbc..cd0e9461e4 100644 --- a/internal/frontend/repl/analysis.go +++ b/internal/frontend/repl/analysis.go @@ -251,8 +251,7 @@ func (s *Session) runAnalysis(inv analysisInvocation) (caseRun, error) { // analysisSymbol resolves the case an invocation names. It is resolved before the // runtime is built, so a misspelling is reported as one whatever the session holds. func (s *Session) analysisSymbol(inv analysisInvocation) (*symbols.Symbol, string, error) { - doc := s.ws.Document(docName) - if doc == nil || doc.Scope == nil { + if !s.hasDeclarations() { return nil, "", errors.New("no declarations loaded") } return s.lookupSymbolOfKinds(inv.name, @@ -295,7 +294,7 @@ func (s *Session) runAnalysisIn(x execution, ctx *runtime.Context, inv analysisI // A usage owned by a type is a feature of an object of that type, which the // session holds when one was created; a package-level case has no such owner. self := nestedCaseOwner(sym, fqn, objects) - runScope := declaringScope(sym, s.ws.Document(docName).Scope) + runScope := declaringScope(sym, s.rootScopeOf(sym)) // A verification case runs the same body; asking the run for its verdict too // reports it beside what the run computed. diff --git a/internal/frontend/repl/filedocs_test.go b/internal/frontend/repl/filedocs_test.go new file mode 100644 index 0000000000..6f4f8fdbdc --- /dev/null +++ b/internal/frontend/repl/filedocs_test.go @@ -0,0 +1,334 @@ +package repl + +import ( + "errors" + "fmt" + "os" + "path/filepath" + "slices" + "sort" + "strings" + "testing" + + "github.com/Open-MBEE/OpenSysML/internal/syntax/source" + "github.com/Open-MBEE/OpenSysML/internal/workspace/model" +) + +// A file loaded from the command line is a document of its own: a root-level +// import in one file surfaces its names in that file's root namespace only, so +// the other files on the command line do not see them, exactly as the editor +// and the workspace report it. +func TestLoadedFilesDoNotShareRootImports(t *testing.T) { + dir := t.TempDir() + paths := []string{ + writeFile(t, filepath.Join(dir, "a.sysml"), "import ScalarValues::*;\npackage A { attribute x : Real; }\n"), + writeFile(t, filepath.Join(dir, "b.sysml"), "package B { attribute y : Real; }\n"), + } + s := NewSession() + if _, err := s.LoadFilesSummary(paths); err != nil { + t.Fatal(err) + } + var inB []string + for _, d := range s.LocatedDiagnostics() { + if d.File == paths[1] { + inB = append(inB, d.Message) + } + } + if len(inB) != 1 || !strings.Contains(inB[0], "unresolved reference: Real") { + t.Errorf("b.sysml should not see a.sysml's root import; its diagnostics: %q", inB) + } + if got, want := cliDiagnostics(t, paths), workspaceDiagnostics(t, paths); strings.Join(got, "\n") != strings.Join(want, "\n") { + t.Errorf("the CLI reported:\n%s\nwant, as the workspace does:\n%s", strings.Join(got, "\n"), strings.Join(want, "\n")) + } +} + +// Two files declaring the same root package are two root namespaces of one name, +// not a duplicate; a reference resolves to the first declaration. +func TestLoadedFilesDeclaringOneRootPackageAreNotDuplicates(t *testing.T) { + dir := t.TempDir() + paths := []string{ + writeFile(t, filepath.Join(dir, "a.sysml"), "package A { part def X; }\n"), + writeFile(t, filepath.Join(dir, "b.sysml"), "package A { part def Y; }\n"), + writeFile(t, filepath.Join(dir, "c.sysml"), "package C { part x : A::X; }\n"), + } + s := NewSession() + out, err := s.LoadFilesSummary(paths) + if err != nil { + t.Fatal(err) + } + if s.HasErrors() { + t.Errorf("the files did not validate clean:\n%s", strings.Join(s.DiagnosticLines(), "\n")) + } + for _, line := range out { + if strings.Contains(line, "Duplicate") { + t.Errorf("a repeated root package was reported as a duplicate: %s", line) + } + } + if got, want := cliDiagnostics(t, paths), workspaceDiagnostics(t, paths); strings.Join(got, "\n") != strings.Join(want, "\n") { + t.Errorf("the CLI reported:\n%s\nwant, as the workspace does:\n%s", strings.Join(got, "\n"), strings.Join(want, "\n")) + } +} + +// Which declaration of a repeated root name a reference reaches follows the +// documents' name order, as the workspace orders them, not the command line. +func TestRepeatedRootPackageResolvesByDocumentNameNotLoadOrder(t *testing.T) { + dir := t.TempDir() + first := writeFile(t, filepath.Join(dir, "first.sysml"), "package A { part def X; }\n") + second := writeFile(t, filepath.Join(dir, "second.sysml"), "package A { part def Y; }\n") + useX := writeFile(t, filepath.Join(dir, "use-x.sysml"), "package C { part x : A::X; }\n") + useY := writeFile(t, filepath.Join(dir, "use-y.sysml"), "package D { part y : A::Y; }\n") + + for _, paths := range [][]string{{first, second, useX, useY}, {second, first, useY, useX}} { + got := cliDiagnostics(t, paths) + if len(got) != 1 || !strings.Contains(got[0], "use-y.sysml") || !strings.Contains(got[0], "A::Y") { + t.Errorf("loading %v reported:\n%s\nwant only A::Y unresolved: first.sysml's A sorts first", basenames(paths), strings.Join(got, "\n")) + } + if want := workspaceDiagnostics(t, paths); strings.Join(got, "\n") != strings.Join(want, "\n") { + t.Errorf("the CLI reported:\n%s\nwant, as the workspace does:\n%s", strings.Join(got, "\n"), strings.Join(want, "\n")) + } + } +} + +func basenames(paths []string) []string { + out := make([]string, len(paths)) + for i, p := range paths { + out[i] = filepath.Base(p) + } + return out +} + +// The prompt's transcript is a document of its own too: a root-level import in +// a loaded file does not serve what is typed after %load, though the file's +// root packages are reachable through the global namespace as before. +func TestPromptDoesNotSeeALoadedFilesRootImports(t *testing.T) { + s := NewSession() + path := tempFile(t, "a.sysml", "import ScalarValues::*;\npackage A { attribute x : Real; }\n") + if _, _, err := s.runMeta("%load " + path); err != nil { + t.Fatal(err) + } + res := s.Submit("package P { attribute y : Real; part a : A; }") + var messages []string + for _, d := range res.Diagnostics { + if res.mine(d.Span) { + messages = append(messages, d.Message) + } + } + if len(messages) != 1 || !strings.Contains(messages[0], "unresolved reference: Real") { + t.Errorf("the prompt should resolve A but not the file's import of Real; it reported %q", messages) + } +} + +// An error in a loaded file gates the deeper checks of that file only: a clean +// prompt submission is fully analyzed in its own document, so no blocker is named +// for it, and a later loaded file is not blocked by what the prompt holds either. +func TestLoadedFileErrorsDoNotBlockOtherDocuments(t *testing.T) { + s := NewSession() + bad := tempFile(t, "bad.sysml", "package Bad { part a : Missing; }\n") + if _, _, err := s.runMeta("%load " + bad); err != nil { + t.Fatal(err) + } + res := s.Submit("package Clean { part def A; }") + if note := res.Blocked.note(); note != "" { + t.Errorf("a loaded file's error should not block the prompt's document: %s", note) + } + + s.Submit("package Typed { part b : Absent::B; }") + good := tempFile(t, "good.sysml", "package Good { part def B; }\n") + if note := s.SubmitFiles([]SourceFile{{Name: good, Text: "package Good { part def B; }\n"}}).Blocked.note(); note != "" { + t.Errorf("the prompt's error should not block a loaded file's document: %s", note) + } + // Within the transcript, an earlier typed error still gates the deeper checks, + // and is named once: a load in between does not make it worth saying again. + if s.Submit("package Also { part def C; }").Blocked.note() == "" { + t.Error("the typed unresolved reference should still be named as blocking the prompt") + } + s.SubmitFiles([]SourceFile{{Name: good, Text: "package Good { part def B; }\n"}}) + if note := s.Submit("package More { part def D; }").Blocked.note(); note != "" { + t.Errorf("the standing error was named already; a load does not renew it: %s", note) + } + // A load that resolves the standing error ends its interval: should a reload + // bring the error back, the next prompt is told again. + s.SubmitFiles([]SourceFile{{Name: good, Text: "package Absent { part def B; }\n"}}) + s.SubmitFiles([]SourceFile{{Name: good, Text: "package Good { part def B; }\n"}}) + if s.Submit("package Yet { part def E; }").Blocked.note() == "" { + t.Error("an error resolved by a load and brought back by a reload should be named again") + } +} + +// Every multi-file directory of the fixtures and of the OMG corpora reports the +// same diagnostics loaded from the command line as opened in a workspace. +func TestCommandLineLoadMatchesWorkspace(t *testing.T) { + roots := []struct { + dir string + require string // set in CI, where an absent corpus fails instead of skipping + fetch string + }{ + {dir: "../../../tests/testdata"}, + {dir: "../../../examples"}, + { + dir: "../../../examples/sysml-v2-training", + require: "OPENSYSML_REQUIRE_TRAINING_CORPUS", + fetch: "./scripts/download-training-examples.sh", + }, + { + dir: "../../../examples/pilot-corpora/kerml-examples", + require: "OPENSYSML_REQUIRE_PILOT_CORPORA", + fetch: "./scripts/download-pilot-corpora.sh", + }, + { + dir: "../../../examples/pilot-corpora/sysml-examples", + require: "OPENSYSML_REQUIRE_PILOT_CORPORA", + fetch: "./scripts/download-pilot-corpora.sh", + }, + { + dir: "../../../examples/pilot-corpora/sysml-validation", + require: "OPENSYSML_REQUIRE_PILOT_CORPORA", + fetch: "./scripts/download-pilot-corpora.sh", + }, + } + seen := map[string]bool{} + for _, root := range roots { + if _, err := os.Stat(root.dir); os.IsNotExist(err) { + if os.Getenv(root.require) != "" { + t.Fatalf("%s is missing and %s is set; fetch it with %s", root.dir, root.require, root.fetch) + } + t.Logf("%s is absent (fetch it with %s), so this run proves nothing about it", root.dir, root.fetch) + continue + } + for _, files := range modelDirectories(t, root.dir) { + dir := filepath.Dir(files[0]) + if seen[dir] { + continue + } + seen[dir] = true + t.Run(filepath.ToSlash(dir), func(t *testing.T) { + got, want := cliDiagnostics(t, files), workspaceDiagnostics(t, files) + if strings.Join(got, "\n") != strings.Join(want, "\n") { + t.Errorf("the CLI reported:\n%s\nwant, as the workspace does:\n%s", strings.Join(got, "\n"), strings.Join(want, "\n")) + } + }) + } + } +} + +// modelDirectories walks root and returns the model files of every directory +// holding more than one, each directory's files sorted, the corpora's directories +// under a root that contains them included. +func modelDirectories(t *testing.T, root string) [][]string { + t.Helper() + byDir := map[string][]string{} + err := filepath.WalkDir(root, func(path string, entry os.DirEntry, err error) error { + if err != nil { + return err + } + if entry.IsDir() || source.KindOf(path) == source.KindUnknown { + return nil + } + byDir[filepath.Dir(path)] = append(byDir[filepath.Dir(path)], path) + return nil + }) + if err != nil { + t.Fatalf("scan %s: %v", root, err) + } + dirs := make([]string, 0, len(byDir)) + for dir, files := range byDir { + if len(files) > 1 { + dirs = append(dirs, dir) + } + } + sort.Strings(dirs) + out := make([][]string, 0, len(dirs)) + for _, dir := range dirs { + files := byDir[dir] + sort.Strings(files) + out = append(out, files) + } + return out +} + +// cliDiagnostics loads the files as the command line does and returns what it +// reports, one sorted line per diagnostic. +func cliDiagnostics(t *testing.T, paths []string) []string { + t.Helper() + s := NewSession() + if _, err := s.LoadFilesSummary(paths); err != nil { + t.Fatal(err) + } + var out []string + for _, d := range s.LocatedDiagnostics() { + out = append(out, fmt.Sprintf("%s:%d:%d: %s: %s [%s]", filepath.Base(d.File), d.Line, d.Column, d.Severity, d.Message, d.Code)) + } + sort.Strings(out) + return out +} + +// workspaceDiagnostics opens the files in one workspace, as the editor and the +// corpus gates do, and returns what it reports in the same form as cliDiagnostics. +func workspaceDiagnostics(t *testing.T, paths []string) []string { + t.Helper() + ws := model.NewWorkspace() + contents := make(map[string][]byte, len(paths)) + for _, path := range paths { + content, err := os.ReadFile(path) + if err != nil { + t.Fatal(err) + } + contents[path] = content + ws.Open(path, content, 1) + } + var out []string + for _, path := range paths { + lines := source.New(path, contents[path]).Lines() + for _, d := range ws.Diagnostics(path) { + p := lines.PosAt(d.Span.Offset) + out = append(out, fmt.Sprintf("%s:%d:%d: %s: %s [%s]", filepath.Base(path), p.Line, p.Col, d.Severity, d.Message, d.Code)) + } + } + sort.Strings(out) + return out +} + +// The transcript is kept under one workspace name, so a file of that name is +// refused at the load rather than sharing the document with the typed text. +func TestLoadRefusesAFileNamedAsTheTranscript(t *testing.T) { + dir := t.TempDir() + writeFile(t, filepath.Join(dir, docName), "package FromFile { part def X; }\n") + t.Chdir(dir) + + s := NewSession() + s.Submit("package Typed { part def T; }") + before := s.Text() + + _, err := s.LoadFilesSummary([]string{docName}) + var reserved *ReservedNameError + if !errors.As(err, &reserved) || reserved.Name != docName { + t.Fatalf("LoadFilesSummary(%q) error = %v, want a *ReservedNameError naming it", docName, err) + } + if _, _, err := s.runMeta("%load " + docName); err == nil || !strings.Contains(err.Error(), "reserved") { + t.Fatalf("%%load %s error = %v, want the name refused as reserved", docName, err) + } + if _, err := s.LoadFile(docName); !errors.As(err, &reserved) { + t.Errorf("LoadFile(%q) error = %v, want a *ReservedNameError", docName, err) + } + + // A direct submission has no error to return, so the whole of it is refused + // in the result: nothing accepted, the refusal all it renders. + res := s.SubmitFiles([]SourceFile{ + {Name: "ok.sysml", Text: "package Direct { part def D; }\n"}, + {Name: docName, Text: "package FromFile { part x : Missing; }\n"}, + }) + if !errors.As(res.Refused, &reserved) || len(res.Declared) != 0 || len(res.Origins) != 0 { + t.Errorf("a refused SubmitFiles = {Refused: %v, Declared: %v, Origins: %v}, want a *ReservedNameError and nothing else", + res.Refused, res.Declared, res.Origins) + } + want := []string{"error: cannot load : the name is reserved for the text typed at the prompt"} + if got := renderResult(res, VerbosityNormal); !slices.Equal(got, want) { + t.Errorf("a refused SubmitFiles rendered %q, want %q", got, want) + } + if got := s.Text(); got != before { + t.Errorf("the refused load changed the transcript:\n%s\nwas:\n%s", got, before) + } + if got := strings.Join(s.List(), "\n"); strings.Contains(got, "FromFile") || strings.Contains(got, "Direct") || !strings.Contains(got, "Typed") { + t.Errorf("the refused file's declarations must not enter the session; got %v", s.List()) + } +} diff --git a/internal/frontend/repl/jobs.go b/internal/frontend/repl/jobs.go index 0abbd318fe..38c8c43366 100644 --- a/internal/frontend/repl/jobs.go +++ b/internal/frontend/repl/jobs.go @@ -12,18 +12,27 @@ func (s *Session) Jobs() int { return s.jobs } -// SetJobs sets how many runs of one plan go concurrently from here on. The held -// context, its objects and the debuggers keep going: jobs bound a plan's runs, -// not the session's context. A value below one is a typed error. +// SetJobs sets how many runs of one plan go concurrently from here on, and how +// many files of one load are parsed and analyzed at once. The held context, its +// objects and the debuggers keep going: jobs bound a plan's runs, not the +// session's context. A value below one is a typed error. func (s *Session) SetJobs(jobs int) error { if jobs <= 0 { return &analysis.JobsError{Source: "%jobs", Value: fmt.Sprint(jobs)} } defer s.enter()() - s.jobs = jobs + s.setJobs(jobs) return nil } +// setJobs records a positive job count and sizes the workspace's pool to match. +func (s *Session) setJobs(jobs int) { + s.jobs = jobs + if err := s.ws.SetWorkers(jobs); err != nil { + panic(err) // unreachable: jobs is positive + } +} + // doJobs shows the jobs setting, or sets it when a count is given. func (s *Session) doJobs(args []string) []string { if len(args) > 0 { @@ -31,7 +40,7 @@ func (s *Session) doJobs(args []string) []string { if err != nil { return []string{errPrefix + err.Error()} } - s.jobs = jobs + s.setJobs(jobs) } return []string{fmt.Sprintf("jobs: %d", s.jobs)} } diff --git a/internal/frontend/repl/load.go b/internal/frontend/repl/load.go index 180770e6aa..888ccd116b 100644 --- a/internal/frontend/repl/load.go +++ b/internal/frontend/repl/load.go @@ -53,22 +53,15 @@ func (s *Session) loadPathsReport(paths []string) (LoadReport, error) { if err != nil { return LoadReport{}, err } - files = s.withDependencies(files) - srcs := make([]SourceFile, 0, len(files)) - names := make([]string, 0, len(files)) - for _, file := range files { - name, data, err := project.ReadFile(file) - if err != nil { - return LoadReport{}, readError(name, err) - } - names = append(names, name) - srcs = append(srcs, SourceFile{Name: name, Text: string(data)}) + srcs, err := s.readSources(s.withDependencies(files)) + if err != nil { + return LoadReport{}, err } var loaded []string - if len(files) > 1 { - loaded = append(loaded, fmt.Sprintf("loaded %d files:", len(files))) - for _, name := range names { - loaded = append(loaded, " "+name) + if len(srcs) > 1 { + loaded = append(loaded, fmt.Sprintf("loaded %d files:", len(srcs))) + for _, src := range srcs { + loaded = append(loaded, " "+src.Name) } } found, declared := renderSplit(s.submitFiles(srcs), s.verbosity) @@ -104,6 +97,42 @@ func expandHomes(paths []string) []string { return out } +// readSources reads every path into the file a load submits, under the name it +// is reported by; the error is a *ReadError or a *ReservedNameError. +func (s *Session) readSources(paths []string) ([]SourceFile, error) { + files := make([]SourceFile, 0, len(paths)) + for _, path := range paths { + name, data, err := project.ReadFile(path) + if err != nil { + return nil, readError(name, err) + } + if err := reservedName(name); err != nil { + return nil, err + } + files = append(files, SourceFile{Name: name, Text: string(data)}) + } + return files, nil +} + +// ReservedNameError is a file a load refused because its name is the one the +// session keeps its typed text under, which a loaded file cannot share. +type ReservedNameError struct { + Name string +} + +func (e *ReservedNameError) Error() string { + return fmt.Sprintf("cannot load %s: the name is reserved for the text typed at the prompt", e.Name) +} + +// reservedName is the *ReservedNameError refusing a file named as the +// transcript, nil for any other name. +func reservedName(name string) error { + if name != docName { + return nil + } + return &ReservedNameError{Name: name} +} + // ReadError is a file a load could not read, under the name it is reported by. type ReadError struct { Path string diff --git a/internal/frontend/repl/meta.go b/internal/frontend/repl/meta.go index d3ed99e6a6..e437e9e09b 100644 --- a/internal/frontend/repl/meta.go +++ b/internal/frontend/repl/meta.go @@ -5,7 +5,6 @@ import ( "fmt" "math" "slices" - "sort" "strconv" "strings" "unicode" @@ -796,7 +795,7 @@ func (s *Session) evalIn(name, expr string) ([]string, error) { // contextScope is the namespace a pinned context evaluates in: the element's own // scope, so its members are named without qualification, else the scope it was -// declared in, searched through both session documents. +// declared in, searched through every session document. func (s *Session) contextScope(sym *symbols.Symbol) *symbols.Scope { if sym == nil { return nil @@ -850,14 +849,14 @@ func (s *Session) evalExpr(expr string) ([]string, error) { return literalResult, litErr } - doc := s.ws.Document(docName) + declared := s.hasDeclarations() // The library is indexed with or without session declarations, so a name it // declares is answered from it; only compound expressions, handled below, - // need the session's own document. + // need the session's own documents. ctx, err := s.getOrCreateRuntime() if err != nil { - if doc == nil || doc.Scope == nil { + if !declared { return nil, s.errWithoutDeclarations(expr) } return nil, err @@ -951,12 +950,14 @@ func (s *Session) evalExpr(expr string) ([]string, error) { // A compound expression is evaluated in the session's own namespace; an empty // session has none, so only the library answers there. - if doc == nil || doc.Scope == nil { + if !declared { return s.evalWithoutDeclarations(ctx, expr) } - // Complex expression with feature refs - inject into session context - tempSrc := s.joined() + fmt.Sprintf("\nattribute __eval__ = %s;", expr) + // Complex expression with feature refs - parsed after the transcript, the + // loaded files masked out of it as they are out of the transcript document + typed, _ := s.transcript() + tempSrc := typed + fmt.Sprintf("\nattribute __eval__ = %s;", expr) p := parser.New(source.New("eval", []byte(tempSrc))) root := p.ParseFile() @@ -1853,8 +1854,7 @@ func (s *Session) evalCalc(calcName, argText string) ([]string, []NamedValue, *a // calcSymbol resolves the calc %calc names. It is resolved before the runtime is // built, so a misspelling is reported as one whatever the session holds. func (s *Session) calcSymbol(calcName string) (*symbols.Symbol, error) { - doc := s.ws.Document(docName) - if doc == nil || doc.Scope == nil { + if !s.hasDeclarations() { return nil, errors.New("no declarations loaded") } sym, _, lerr := s.lookupSymbolOfKinds(calcName, symbols.SymbolCalcDef, symbols.SymbolCalcUsage) @@ -2161,31 +2161,21 @@ func (s *Session) doConstraint(name string) ([]string, bool, error) { // promptScope is the namespace a prompt expression is evaluated in: the last // namespace the session declared, whose imports are then visible to it exactly // as they are to a member written there (KerML 8.2.3.5.3). A session that -// declared no namespace evaluates at the document root. Both session documents -// are read, in buffer order, so a namespace loaded from a .kerml file counts. +// declared no namespace evaluates at the document root. Every session document +// is read, in buffer order, so a namespace a loaded file declares counts. func (s *Session) promptScope() *symbols.Scope { docs := s.sessionDocs() if len(docs) == 0 { return nil } - type entry struct { - member ast.Node - scope *symbols.Scope - } - var members []entry - for _, doc := range docs { - if doc.AST == nil || doc.Scope == nil { - continue - } - for _, m := range doc.AST.Members { - members = append(members, entry{m, doc.Scope}) + var members []Member + for _, m := range s.sessionMembers() { + if m.scope != nil { + members = append(members, m) } } - sort.SliceStable(members, func(i, j int) bool { - return members[i].member.Span().Offset < members[j].member.Span().Offset - }) for i := len(members) - 1; i >= 0; i-- { - member := members[i].member + member := members[i].Node if mem, ok := member.(*ast.Membership); ok { member = mem.Member } @@ -2207,7 +2197,7 @@ func (s *Session) promptScope() *symbols.Scope { } } // No namespace to work in: the root holding the last declaration, so a - // top-level member loaded from a .kerml file is still in reach. + // top-level member of the last loaded file is still in reach. if len(members) > 0 { return members[len(members)-1].scope } diff --git a/internal/frontend/repl/montecarlo.go b/internal/frontend/repl/montecarlo.go index 0e8232ac5b..98fda6de36 100644 --- a/internal/frontend/repl/montecarlo.go +++ b/internal/frontend/repl/montecarlo.go @@ -180,8 +180,7 @@ func (m *monteCarloRuns) conclusion() (runtime.AnalysisResult, error) { // monteCarloSample makes count runs of the invocation, each in its own context on objects // made from their declarations; the plan is returned beside a refusal made after an engine ran. func (s *Session) monteCarloSample(inv analysisInvocation, count int64, seed *uint64) (*monteCarloRuns, *analysis.Plan, error) { - doc := s.ws.Document(docName) - if doc == nil || doc.Scope == nil { + if !s.hasDeclarations() { return nil, nil, errors.New("no declarations loaded") } if _, replaying := s.drivenSchedule().Replay(); replaying { @@ -228,7 +227,7 @@ func (s *Session) monteCarloSample(inv analysisInvocation, count int64, seed *ui return nil, nil, err } } - runScope := declaringScope(sym, doc.Scope) + runScope := declaringScope(sym, s.rootScopeOf(sym)) if seed == nil && s.modelSeed.set { session := s.modelSeed.value diff --git a/internal/frontend/repl/openinput_test.go b/internal/frontend/repl/openinput_test.go index 1c2f6dcf23..cfa1135542 100644 --- a/internal/frontend/repl/openinput_test.go +++ b/internal/frontend/repl/openinput_test.go @@ -249,6 +249,35 @@ func TestReloadingAFixedFileClearsItsSyntaxError(t *testing.T) { } } +// The reverse: a file reloaded with its enclosure left open is masked, and takes +// its declarations out of the index with it, so a qualified lookup no longer +// finds what the session no longer holds; the library stays reachable. +func TestReloadingAFileLeftOpenDropsItsSymbols(t *testing.T) { + s := NewSession() + path := filepath.Join(t.TempDir(), "model.sysml") + if err := os.WriteFile(path, []byte("package P { attribute x = 1; }\n"), 0o644); err != nil { + t.Fatal(err) + } + if _, err := s.LoadFile(path); err != nil { + t.Fatal(err) + } + if _, _, err := s.lookupSymbol("P::x"); err != nil { + t.Fatalf("P::x did not resolve after the load: %v", err) + } + if err := os.WriteFile(path, []byte("package P {\n"), 0o644); err != nil { + t.Fatal(err) + } + if _, err := s.LoadFile(path); err != nil { + t.Fatal(err) + } + if sym, _, err := s.lookupSymbol("P::x"); err == nil { + t.Errorf("P::x should be gone with the file that declared it, found %v", sym) + } + if _, _, err := s.lookupSymbol("ScalarValues::Real"); err != nil { + t.Errorf("the library should still answer a qualified lookup: %v", err) + } +} + // Typed input that leaves an enclosure open is masked the same way, and the // declarations already in the buffer are untouched. func TestOpenTypedSubmissionKeepsTheBuffer(t *testing.T) { diff --git a/internal/frontend/repl/print.go b/internal/frontend/repl/print.go index edd7ff63c7..83e3b66336 100644 --- a/internal/frontend/repl/print.go +++ b/internal/frontend/repl/print.go @@ -54,7 +54,7 @@ func (s *Session) printElement(name string) ([]string, bool, error) { shown = name } var doc *model.Document - if sym != nil && (sym.DocName == docName || sym.DocName == kermlDocName) { + if sym != nil { doc = s.ws.Document(sym.DocName) } if doc == nil || sym == nil || sym.Decl == nil { diff --git a/internal/frontend/repl/render.go b/internal/frontend/repl/render.go index 5b0fc385c9..d9bcfad1b8 100644 --- a/internal/frontend/repl/render.go +++ b/internal/frontend/repl/render.go @@ -8,16 +8,17 @@ import ( "golang.org/x/text/width" + "github.com/Open-MBEE/OpenSysML/internal/semantic/symbols" "github.com/Open-MBEE/OpenSysML/internal/syntax/ast" "github.com/Open-MBEE/OpenSysML/internal/syntax/diag" "github.com/Open-MBEE/OpenSysML/internal/syntax/source" ) -// Result is the outcome of one Submit: the top-level members parsed from the -// accumulated buffer (for the success summary), the names this submission -// declared, and any analysis diagnostics over the whole document. +// Result is the outcome of one Submit: the top-level members of the session's +// documents (for the success summary), the names this submission declared, and +// any analysis diagnostics over the whole buffer. type Result struct { - Members []ast.Node // top-level members of the AST (Task 5 renders these) + Members []Member // top-level members of the session documents, in buffer order Declared []string // names introduced by THIS submission Diagnostics []diag.Diagnostic // eager analysis over the whole buffer Source string // the full joined content (Task 6 caret rendering) @@ -25,6 +26,10 @@ type Result struct { Origins []Origin // the files of THIS submission, in buffer order Notices []string // side effects of the submission, e.g. a debugging session it ended + // Refused is the *ReservedNameError a submission was refused for, nil when + // it was accepted; a refused submission changed nothing. + Refused error + // Blocked names the unresolved error that stopped the deeper checks from // running over this submission, nil when they ran or when the session already // reported that error. @@ -38,6 +43,19 @@ type Result struct { // masked locates the submissions kept out of the analyzed buffer, whose // findings gated no validation tier. masked []source.Span + + // foreign locates the snippets analyzed in another document than this + // submission's, whose findings gated none of its validation tiers. + foreign []source.Span +} + +// Member is one top-level member of a session document; Offset is where the +// member begins in the buffer, a loaded file's document having offsets of its own. +type Member struct { + Node ast.Node + Offset int + // scope is the root scope of the document declaring the member. + scope *symbols.Scope } // Origin locates one file of a submission in the buffer, so a diagnostic is @@ -110,10 +128,10 @@ func (r Result) holdsMine(span source.Span) bool { } // renderSummary returns one accepted line per top-level member: "✓ ". -func renderSummary(members []ast.Node) []string { +func renderSummary(members []Member) []string { out := make([]string, 0, len(members)) for _, m := range members { - if line := renderMember(m); line != "" { + if line := renderMember(m.Node); line != "" { out = append(out, "✓ "+line) } } @@ -329,6 +347,9 @@ func renderResult(r Result, v Verbosity) []string { // analysis found apart from what the submission declared, so a caller outside // the prompt can send the two to different streams. func renderSplit(r Result, v Verbosity) (found, declared []string) { + if r.Refused != nil { + return []string{"error: " + r.Refused.Error()}, nil + } if v >= VerbosityDebug { // Everything the analysis produced over the whole buffer, at // buffer-absolute positions, plus where this submission landed in it. @@ -355,6 +376,9 @@ func renderSplit(r Result, v Verbosity) (found, declared []string) { // text just read rather than about the analysis of the model as a whole: a load // that defers the analysis still says why a file could not be read. func renderSyntax(r Result, v Verbosity) []string { + if r.Refused != nil { + return []string{"error: " + r.Refused.Error()} + } // A finding about the notation is no reason a file could not be read, and the // analysis this load defers reports it, so reporting it here would report it twice. var diags []diag.Diagnostic @@ -427,13 +451,18 @@ func (b *blocker) note() string { // blockedBy reports the unresolved error that stopped the deeper checks from // running over this submission: a standing error is named on the first -// submission whose report says so, not on every one after it. -func (s *Session) blockedBy(r Result) *blocker { +// submission whose report says so, not on every one after it. A load shares no +// document with the transcript, so it names nothing; one that resolves the +// standing error lets it be named again should it return. +func (s *Session) blockedBy(r Result, load bool) *blocker { b := r.analysisBlocked() if b == nil { s.notedBlocker.record("") return nil } + if load { + return nil + } key := b.key() if key == s.notedBlocker.reportedKey() { return nil @@ -467,7 +496,7 @@ func (n *blockerNote) record(key string) { func (r Result) analysisBlocked() *blocker { var first *blocker for _, d := range r.Diagnostics { - if !d.Blocking() || r.mine(d.Span) || r.isMasked(d.Span) { + if !d.Blocking() || r.mine(d.Span) || covers(r.masked, d.Span) || covers(r.foreign, d.Span) { continue } if first != nil { @@ -479,10 +508,10 @@ func (r Result) analysisBlocked() *blocker { return first } -// isMasked reports whether a span falls in a submission that was kept out of -// the analyzed buffer, so its errors blocked nothing. -func (r Result) isMasked(span source.Span) bool { - for _, m := range r.masked { +// covers reports whether a span starts in one of the snippets located, whose +// errors blocked nothing of the submission's. +func covers(snippets []source.Span, span source.Span) bool { + for _, m := range snippets { // End() included: a submission that does not close its own text is // reported at its end as often as inside it. if span.Offset >= m.Offset && span.Offset <= m.End() { @@ -505,13 +534,13 @@ func hasError(diags []diag.Diagnostic) bool { // within narrows the result to one span of the submission — one file of a load // of several — so what is reported as its own is scoped to that text alone. A -// member is the file's when it begins there: the last member of a document runs -// on over the other language's text masked out after it. +// member is the file's when it begins there: the transcript's last member runs +// on over the files' text masked out after it. func (r Result) within(span source.Span) Result { r.own = []source.Span{span} - members := make([]ast.Node, 0, len(r.Members)) + members := make([]Member, 0, len(r.Members)) for _, m := range r.Members { - if at := m.Span().Offset; at >= span.Offset && at < span.End() { + if m.Offset >= span.Offset && m.Offset < span.End() { members = append(members, m) } } @@ -521,10 +550,10 @@ func (r Result) within(span source.Span) Result { // ownMembers returns the top-level members this submission contributed, so a // summary does not re-announce everything typed earlier in the session. -func (r Result) ownMembers() []ast.Node { - out := make([]ast.Node, 0, len(r.Members)) +func (r Result) ownMembers() []Member { + out := make([]Member, 0, len(r.Members)) for _, m := range r.Members { - if r.holdsMine(m.Span()) { + if r.holdsMine(source.Span{Offset: m.Offset, Len: m.Node.Span().Len}) { out = append(out, m) } } diff --git a/internal/frontend/repl/run.go b/internal/frontend/repl/run.go index d7c7484c40..0d8ad3a341 100644 --- a/internal/frontend/repl/run.go +++ b/internal/frontend/repl/run.go @@ -11,7 +11,6 @@ import ( "github.com/Open-MBEE/OpenSysML/internal/semantic/semantics" "github.com/Open-MBEE/OpenSysML/internal/syntax/diag" "github.com/Open-MBEE/OpenSysML/internal/syntax/source" - "github.com/Open-MBEE/OpenSysML/internal/workspace/project" ) // errRuntimeInit marks a runtime the session could not create at all, which the @@ -54,17 +53,13 @@ func errorLines(lines []string, _ []NamedValue, err error) ([]string, bool, erro // LoadFile submits path and the files beside and below it that declare a root // namespace it imports as one submission, returning the lines `%load` prints. -// A lone "-" reads standard input; the error is a file it could not read. +// A lone "-" reads standard input; the error is a file it could not read or +// one named as the transcript is. func (s *Session) LoadFile(path string) ([]string, error) { defer s.enter()() - paths := s.withDependencies([]string{expandHome(path)}) - files := make([]SourceFile, 0, len(paths)) - for _, p := range paths { - name, data, err := project.ReadFile(p) - if err != nil { - return nil, readError(name, err) - } - files = append(files, SourceFile{Name: name, Text: string(data)}) + files, err := s.readSources(s.withDependencies([]string{expandHome(path)})) + if err != nil { + return nil, err } return renderResult(s.submitFiles(files), s.verbosity), nil } @@ -77,19 +72,16 @@ func (s *Session) LoadFileSummary(path string) ([]string, error) { return s.LoadFilesSummary([]string{path}) } -// LoadFilesSummary is LoadFileSummary over every path as one submission, indexed and -// analyzed once, each file still summarized on its own; a read failure is a *ReadError. -// Files beside and below the paths that declare an imported root namespace load too. +// LoadFilesSummary is LoadFileSummary over every path as one submission, each +// file a document of its own, indexed together and each summarized on its own; +// a read failure is a *ReadError and a file named as the transcript is a +// *ReservedNameError. Files beside and below the paths that declare an imported +// root namespace load too. func (s *Session) LoadFilesSummary(paths []string) ([]string, error) { defer s.enter()() - paths = s.withDependencies(expandHomes(paths)) - files := make([]SourceFile, 0, len(paths)) - for _, path := range paths { - name, data, err := project.ReadFile(path) - if err != nil { - return nil, readError(name, err) - } - files = append(files, SourceFile{Name: name, Text: string(data)}) + files, err := s.readSources(s.withDependencies(expandHomes(paths))) + if err != nil { + return nil, err } res, byFile, whole := s.submitEach(files) var lines []string diff --git a/internal/frontend/repl/session.go b/internal/frontend/repl/session.go index 0849f7beac..895586ec97 100644 --- a/internal/frontend/repl/session.go +++ b/internal/frontend/repl/session.go @@ -27,23 +27,17 @@ import ( "github.com/Open-MBEE/OpenSysML/internal/workspace/model" ) -// docName is the in-memory workspace key for the accumulated REPL buffer. -// Text loaded from a .kerml file keeps that file's language: it is masked out -// of docName and analyzed in kermlDocName, whose name carries the KerML kind -// the parser's file-kind gates read. Both documents span the same joined -// buffer byte for byte, so every offset locates the same snippet in either. +// docName is the workspace key of the transcript: the typed submissions, joined, +// with each loaded file masked out — a file is a workspace document of its own. const docName = "" -// kermlDocName is the workspace key for the buffer's KerML text. -const kermlDocName = ".kerml" - // parseDocName is the document a snippet from origin is parsed and analyzed -// in, which carries the kind of the file it was loaded from. +// in: the file itself when it was loaded from one, else the transcript. func parseDocName(origin string) string { - if source.KindOf(origin) == source.KindKerML { - return kermlDocName + if origin == "" { + return docName } - return docName + return origin } // snippet is one accepted submission source, the top-level names it declares, @@ -75,7 +69,8 @@ type snippet struct { diags []diag.Diagnostic } -// Session accumulates submissions into a single implicit document. +// Session accumulates submissions: what is typed into the transcript document, +// and each loaded file into a document of its own. type Session struct { // mu serializes commands; state guards the session for readers beside one // (Complete answers Tab while a line evaluates). Exported commands take both, @@ -96,9 +91,10 @@ type Session struct { // replaced is a context a debugging session still runs against, whose identity // sequence the context built next takes over. replaced *runtime.Context - idx *symbols.Index // index over the session document, shared by lookup and runtime + idx *symbols.Index // index over the session documents, shared by lookup and runtime libSource libs.Source // the library files idx holds, for their spans' text - idxVersion int // document version idx holds, 0 when it holds none + idxVersion int // session version idx holds, 0 when it holds none + idxDocs []string // the session documents idx holds, taken back when they go about *semantics.AboutIndex // the `about` annotations of idx, renewed with it and shared by every runtime model over it names *nameTable // simple names of the documents, rebuilt when their scope trees change instances map[string]*runtime.Instance // FQN -> instance for %instantiate tracking @@ -294,16 +290,17 @@ func (s *stateSession) selfOf() string { // NewSession returns a session over a fresh workspace. func NewSession() *Session { - return &Session{ + s := &Session{ ws: model.NewWorkspace(), instances: make(map[string]*runtime.Instance), budgets: runtime.DefaultBudgets(), - jobs: analysis.DefaultJobs(), engines: engines.Default(), verbosity: VerbosityNormal, toolVersion: "sysml dev", now: time.Now, } + s.setJobs(analysis.DefaultJobs()) + return s } // SetToolVersion names the tool a recorded run's provenance reports. @@ -415,15 +412,35 @@ func (s *Session) accept(origin, src string) { // A loaded file supersedes only itself and what the prompt said about the same // names, since several files of one model commonly open the same package. func (s *Session) acceptFrom(origin, src string) (declared []string, drops []dropReport) { - p := parser.New(source.New(parseDocName(origin), []byte(src))) - root := p.ParseFile() + return s.acceptParsed(origin, src, preparse(origin, src)) +} + +// parsed is what a submission's text parses to, taken before it is accepted so +// the files of one load can be parsed at once. +type parsed struct { + p *parser.Parser + root *ast.RootNamespace + closes bool +} + +// preparse parses src as the submission from origin, and probes whether it +// closes its own text. +func preparse(origin, src string) parsed { + doc := parseDocName(origin) + p := parser.New(source.New(doc, []byte(src))) + return parsed{p: p, root: p.ParseFile(), closes: closesItsOwnText(doc, src)} +} + +// acceptParsed is acceptFrom over a parse already taken. +func (s *Session) acceptParsed(origin, src string, pre parsed) (declared []string, drops []dropReport) { + p, root := pre.p, pre.root names := declaredNames(root) declared = names text := src // A submission that does not close its own text is masked out of the buffer // rather than left to absorb the submissions after it, and declares nothing: // what the parser recovered from it is not what was meant. - if !closesItsOwnText(parseDocName(origin), src) { + if !pre.closes { key := fileKeyOf(origin) if key != "" { // Re-reading the file supersedes what it declared before, which it no @@ -635,7 +652,7 @@ const sessionOrigin = "" // write the session's text as a document of their own. const SessionOrigin = sessionOrigin -// joined is the buffer the session analyzes: every accepted submission, with a +// joined is the buffer the session presents: every accepted submission, with a // submission that does not close its own text masked out so it cannot change how // the others parse. Masking is byte for byte, so every offset still locates the // snippet and line it came from. @@ -651,15 +668,13 @@ func (s *Session) joined() string { return strings.Join(parts, "\n") } -// joinedFor is the buffer one session document analyzes: joined, with the -// snippets of the other language masked out too, so each document parses its -// own snippets as the kind its name carries while keeping every offset. The -// second result reports whether any snippet of that language survives. -func (s *Session) joinedFor(name string) (string, bool) { +// transcript is the buffer the transcript document analyzes: joined, with the +// loaded files masked out too; the second result reports whether any typed text remains. +func (s *Session) transcript() (string, bool) { parts := make([]string, len(s.snippets)) found := false for i, sn := range s.snippets { - if sn.open || parseDocName(sn.origin) != name { + if sn.open || sn.origin != "" { parts[i] = maskedText(sn.src) continue } @@ -669,6 +684,34 @@ func (s *Session) joinedFor(name string) (string, bool) { return strings.Join(parts, "\n"), found } +// openDocuments brings the workspace to the session's documents: the transcript, +// and one document per loaded file that parses, gone when its snippet goes. +func (s *Session) openDocuments() { + var inputs []model.Input + if typed, found := s.transcript(); found { + inputs = append(inputs, model.Input{Name: docName, Content: []byte(typed), Version: s.version}) + } else { + s.ws.Remove(docName) + } + live := make(map[string]bool, len(s.snippets)) + for _, sn := range s.snippets { + if sn.origin == "" || sn.open { + continue + } + live[sn.origin] = true + if doc := s.ws.Document(sn.origin); doc == nil || doc.Version != sn.gen { + inputs = append(inputs, model.Input{Name: sn.origin, Content: []byte(sn.src), Version: sn.gen}) + } + } + for _, name := range s.ws.DocumentNames() { + if name != docName && !live[name] { + s.ws.Remove(name) + } + } + // One batch: the documents are parsed at once and the imports expanded once. + s.ws.OpenAll(inputs) +} + // text is the buffer as it was submitted, masking nothing: what %save writes // back, so work the parser could not read is not lost. func (s *Session) text() string { @@ -720,39 +763,49 @@ func (s *Session) maskedSpans() []source.Span { return out } -// openDiagnostics reports the findings of the masked submissions, located in the -// session buffer so every surface places them in the file they came from. -func (s *Session) openDiagnostics() []diag.Diagnostic { - var out []diag.Diagnostic +// foreignSpans locates the loaded files in the buffer, each a document of its +// own whose findings gated nothing of the transcript's. +func (s *Session) foreignSpans() []source.Span { + var out []source.Span acc := 0 for _, sn := range s.snippets { - if sn.open { - for _, d := range sn.diags { - d.Span.Offset += acc - out = append(out, d) - } + if sn.origin != "" { + out = append(out, source.Span{Offset: acc, Len: len(sn.src)}) } - acc += len(sn.src) + 1 // the newline joined() writes between snippets + acc += len(sn.src) + 1 } return out } -// diagnostics reports the analysis of the buffer together with the syntax errors -// of the submissions masked out of it. The masked text is blanked rather than -// removed, so what the analysis finds is about the submissions that did parse -// and is reported as it stands. Both session documents share the buffer's -// coordinates, so their findings interleave by offset. +// diagnostics reports the analysis of every session document and the syntax errors +// of the masked submissions, each moved to where its text sits in the session buffer. func (s *Session) diagnostics() []diag.Diagnostic { - analyzed := append([]diag.Diagnostic{}, s.ws.Diagnostics(docName)...) - analyzed = append(analyzed, s.ws.Diagnostics(kermlDocName)...) - open := s.openDiagnostics() - if len(open) == 0 { - sort.SliceStable(analyzed, func(i, j int) bool { return analyzed[i].Span.Offset < analyzed[j].Span.Offset }) - return analyzed - } - out := make([]diag.Diagnostic, 0, len(analyzed)+len(open)) - out = append(out, analyzed...) - out = append(out, open...) + names := []string{docName} + for _, sn := range s.snippets { + if sn.origin != "" && !sn.open { + names = append(names, sn.origin) + } + } + // One batch: the documents not analyzed yet are analyzed at once. + analyzed := s.ws.DiagnosticsAll(names) + out := append([]diag.Diagnostic{}, analyzed[0]...) + next := 1 + acc := 0 + for _, sn := range s.snippets { + var own []diag.Diagnostic + switch { + case sn.open: + own = sn.diags + case sn.origin != "": + own = analyzed[next] + next++ + } + for _, d := range own { + d.Span.Offset += acc + out = append(out, d) + } + acc += len(sn.src) + 1 // the newline joined() writes between snippets + } sort.SliceStable(out, func(i, j int) bool { return out[i].Span.Offset < out[j].Span.Offset }) return out } @@ -830,12 +883,34 @@ func (s *Session) submitAll(srcs []string) Result { // SubmitFiles accumulates every file as one submission: all of them are accepted // before the buffer is reindexed and analyzed, so a declaration in one resolves // against the others no matter which order they arrive in. This is what makes -// loading a multi-file project order-independent. +// loading a multi-file project order-independent. A file's Name is its +// workspace document, so a file named as the transcript is refused: nothing is +// accepted and the result carries the *ReservedNameError as Refused. func (s *Session) SubmitFiles(files []SourceFile) Result { defer s.enter()() + for _, f := range files { + if err := reservedName(f.Name); err != nil { + return s.refuse(err) + } + } return s.submitFiles(files) } +// refuse is the result of a submission no part of which was accepted: the +// session as it stands, with nothing of its own but the refusal. +func (s *Session) refuse(err error) Result { + text := s.text() + return Result{ + Members: s.sessionMembers(), + Diagnostics: s.diagnostics(), + Source: text, + Offset: len(text), + Refused: err, + masked: s.maskedSpans(), + foreign: s.foreignSpans(), + } +} + func (s *Session) submitFiles(files []SourceFile) Result { res, _, _ := s.submitEach(files) return res @@ -850,9 +925,14 @@ func (s *Session) submitEach(files []SourceFile) (res Result, byFile [][]string, ) seen := map[string]bool{} s.version++ + load := len(files) > 0 && files[0].Name != "" byFile = make([][]string, len(files)) + parses := make([]parsed, len(files)) + model.ParallelFor(s.jobs, len(files), func(i int) { + parses[i] = preparse(files[i].Name, files[i].Text) + }) for i, f := range files { - names, dropped := s.acceptFrom(f.Name, f.Text) + names, dropped := s.acceptParsed(f.Name, f.Text, parses[i]) for _, name := range names { if !seen[name] { seen[name] = true @@ -887,13 +967,14 @@ func (s *Session) submitEach(files []SourceFile) (res Result, byFile [][]string, Origins: s.origins(), own: own, masked: s.maskedSpans(), + foreign: s.foreignSpans(), Notices: notices, } - res.Blocked = s.blockedBy(res) + res.Blocked = s.blockedBy(res, load) return res, byFile, whole } -// rebuildOver replaces the open document and everything derived from it — the +// rebuildOver replaces the open documents and everything derived from them — the // runtime context, the resolutions held objects and debugging sessions were // made against — after the snippets changed, reporting what it carried over. func (s *Session) rebuildOver(drops []dropReport) []string { @@ -901,14 +982,8 @@ func (s *Session) rebuildOver(drops []dropReport) []string { // before the new text replaces that resolution, so what the new document does // not change can be told apart from what it does. over := s.recordCarryover() - sysml, _ := s.joinedFor(docName) - s.ws.Open(docName, []byte(sysml), s.version) - if kerml, found := s.joinedFor(kermlDocName); found { - s.ws.Open(kermlDocName, []byte(kerml), s.version) - } else { - s.ws.Remove(kermlDocName) - } - // The document is a new AST and scope tree, so the context derived from the + s.openDocuments() + // The documents are new ASTs and scope trees, so the context derived from the // previous one is replaced; the objects it holds are carried into the new one // where the declarations they were materialized against are unchanged. The // index is re-used and brought up to date on the next lookup instead, which is @@ -1156,17 +1231,13 @@ func (s *Session) Clear() []string { // goes is reported and recorded rather than silently emptied. func (s *Session) clear() []string { notices, lost := s.resetLoss() - s.ws.Remove(docName) - s.ws.Remove(kermlDocName) + for _, name := range s.ws.DocumentNames() { + s.ws.Remove(name) + } s.snippets = nil s.version = 0 s.rtCtx, s.replaced = nil, nil - if s.idx != nil { - // Drop the documents, keep the library the index was built with. - s.idx.RemoveDocument(docName) - s.idx.RemoveDocument(kermlDocName) - s.idxVersion, s.about = 0, semantics.NewAboutIndex() - } + s.dropIndexedDocs() s.instances = make(map[string]*runtime.Instance) s.unnamed, s.given = nil, nil s.lost = lost @@ -1269,49 +1340,127 @@ func (s *Session) newRuntimeOver(model *runtime.Model) (*runtime.Context, error) // and the ones its wildcard imports surfaced, so a submission costs its own // document rather than a reload of the library. func (s *Session) symbolIndex() *symbols.Index { - doc := s.ws.Document(docName) - if doc == nil || doc.Scope == nil { + docs := s.sessionDocs() + if !hasScope(docs) { + s.dropIndexedDocs() return nil } if s.idx == nil { s.idx, s.libSource = model.NewIndexWithStdlib() s.about = semantics.NewAboutIndex() - } else if s.idxVersion == doc.Version { + } else if s.idxVersion == s.version { return s.idx } - s.idx.AddDocument(docName, doc.AST) - if kdoc := s.ws.Document(kermlDocName); kdoc != nil { - s.idx.AddDocument(kermlDocName, kdoc.AST) - } else { - s.idx.RemoveDocument(kermlDocName) + live := make(map[string]bool, len(docs)) + for _, doc := range docs { + live[doc.Name] = true + } + for _, name := range s.idxDocs { + if !live[name] { + s.idx.RemoveDocument(name) + } + } + s.idxDocs = s.idxDocs[:0] + for _, doc := range docs { + s.idx.AddDocument(doc.Name, doc.AST) + s.idxDocs = append(s.idxDocs, doc.Name) } s.idx.ExpandWildcardImports() - s.idxVersion, s.about = doc.Version, semantics.NewAboutIndex() + s.idxVersion, s.about = s.version, semantics.NewAboutIndex() return s.idx } -// sessionDocs returns the session's open documents, the SysML buffer first, -// so a caller reading the whole session reads both languages. -func (s *Session) sessionDocs() []*model.Document { - var out []*model.Document - for _, name := range []string{docName, kermlDocName} { - if doc := s.ws.Document(name); doc != nil { - out = append(out, doc) +// dropIndexedDocs takes the session's documents back out of the index, keeping +// the library it was built with. +func (s *Session) dropIndexedDocs() { + if s.idx == nil { + return + } + for _, name := range s.idxDocs { + s.idx.RemoveDocument(name) + } + s.idxDocs = nil + s.idxVersion, s.about = 0, semantics.NewAboutIndex() +} + +// hasScope reports whether any of the documents built a scope tree. +func hasScope(docs []*model.Document) bool { + for _, doc := range docs { + if doc.Scope != nil { + return true } } - return out + return false +} + +// hasDeclarations reports whether the session holds a document with a scope tree. +func (s *Session) hasDeclarations() bool { + return hasScope(s.sessionDocs()) } -// sessionMembers returns the top-level members of both session documents in -// buffer order, which their shared coordinates make the span order. -func (s *Session) sessionMembers() []ast.Node { - var out []ast.Node +// rootScopeOf is the root scope of the session document declaring sym, and for a +// symbol the session declares nowhere that of its first document with one. +func (s *Session) rootScopeOf(sym *symbols.Symbol) *symbols.Scope { + if doc := s.ws.Document(sym.DocName); doc != nil && doc.Scope != nil { + return doc.Scope + } for _, doc := range s.sessionDocs() { - if doc.AST != nil { - out = append(out, doc.AST.Members...) + if doc.Scope != nil { + return doc.Scope + } + } + return nil +} + +// locatedDoc is a session document with the buffer offset its text begins at. +type locatedDoc struct { + doc *model.Document + base int +} + +// locatedDocs returns the transcript, whose offsets are the buffer's, then each +// loaded file's document at the offset its text sits in the buffer. +func (s *Session) locatedDocs() []locatedDoc { + var out []locatedDoc + if doc := s.ws.Document(docName); doc != nil { + out = append(out, locatedDoc{doc: doc}) + } + acc := 0 + for _, sn := range s.snippets { + if sn.origin != "" { + if doc := s.ws.Document(sn.origin); doc != nil { + out = append(out, locatedDoc{doc: doc, base: acc}) + } + } + acc += len(sn.src) + 1 // the newline joined() writes between snippets + } + return out +} + +// sessionDocs returns the session's documents, the transcript first and then the +// loaded files in buffer order. +func (s *Session) sessionDocs() []*model.Document { + located := s.locatedDocs() + out := make([]*model.Document, len(located)) + for i, l := range located { + out[i] = l.doc + } + return out +} + +// sessionMembers returns the top-level members of every session document in +// buffer order, each offset by where its document's text sits. +func (s *Session) sessionMembers() []Member { + var out []Member + for _, l := range s.locatedDocs() { + if l.doc.AST == nil { + continue + } + for _, m := range l.doc.AST.Members { + out = append(out, Member{Node: m, Offset: l.base + m.Span().Offset, scope: l.doc.Scope}) } } - sort.SliceStable(out, func(i, j int) bool { return out[i].Span().Offset < out[j].Span().Offset }) + sort.SliceStable(out, func(i, j int) bool { return out[i].Offset < out[j].Offset }) return out } diff --git a/internal/frontend/repl/sweep.go b/internal/frontend/repl/sweep.go index 64216d108b..71a8c2bbfb 100644 --- a/internal/frontend/repl/sweep.go +++ b/internal/frontend/repl/sweep.go @@ -157,8 +157,7 @@ func sweepLabel(inv analysisInvocation, draws sweepDraws) string { // row per value in a context of its own, held objects made there from their declarations or // from one image of the held graph; the session's state is released while the rows run. func (s *Session) runSweep(inv analysisInvocation, specs []sweepSpec, draws sweepDraws) (runtime.SweepTable, *analysis.Plan, error) { - doc := s.ws.Document(docName) - if doc == nil || doc.Scope == nil { + if !s.hasDeclarations() { return runtime.SweepTable{}, nil, errors.New("no declarations loaded") } sym, fqn, err := s.lookupSymbolOfKinds(inv.name, @@ -213,7 +212,7 @@ func (s *Session) runSweep(inv analysisInvocation, specs []sweepSpec, draws swee if err != nil { return runtime.SweepTable{}, nil, err } - runScope := declaringScope(sym, doc.Scope) + runScope := declaringScope(sym, s.rootScopeOf(sym)) run := func(rt *runtime.Context, bindings []runtime.SweepBinding) (runtime.SweepRunResult, error) { row, err := s.rowObjects(rt, args.objects, image) diff --git a/internal/frontend/repl/view.go b/internal/frontend/repl/view.go index 6171947730..c0a71b35b9 100644 --- a/internal/frontend/repl/view.go +++ b/internal/frontend/repl/view.go @@ -224,14 +224,16 @@ func (s *Session) Views() ([]model.ViewInfo, error) { func (s *Session) symbolsInLoadOrder(in func(*symbols.Scope) []*symbols.Symbol) []*symbols.Symbol { idx := s.browseIndex() var out []*symbols.Symbol - for _, doc := range s.sessionDocs() { - out = append(out, in(idx.DocumentRoot(doc.Name))...) - } - // The language documents are masked copies of one joined buffer, so their - // spans share coordinates and sorting restores submission order. - sort.SliceStable(out, func(i, j int) bool { - return out[i].DeclSpan.Offset < out[j].DeclSpan.Offset - }) + // Each document's symbols are placed where its text sits in the buffer, so + // sorting restores submission order across the documents. + at := make(map[*symbols.Symbol]int) + for _, l := range s.locatedDocs() { + for _, sym := range in(idx.DocumentRoot(l.doc.Name)) { + at[sym] = l.base + sym.DeclSpan.Offset + out = append(out, sym) + } + } + sort.SliceStable(out, func(i, j int) bool { return at[out[i]] < at[out[j]] }) return out } @@ -249,10 +251,10 @@ func (s *Session) viewRenderer() (*view.Renderer, error) { return view.NewRenderer(model, resolver, s.sessionSourceText()), nil } -// sessionSourceFile locates the file a span of the session buffer was loaded from, so a -// location a declaration states relative to its file resolves against that file, not the buffer. +// sessionSourceFile locates the file a span of a session document was loaded from: +// a loaded file is a document named for its path; the transcript's spans are typed. func (s *Session) sessionSourceFile(doc string, span source.Span) string { - if doc != docName && doc != kermlDocName { + if doc != docName { return source.FileNamed(doc, span) } sn, _ := s.snippetAt(span.Offset) diff --git a/internal/frontend/usage/environment.go b/internal/frontend/usage/environment.go index baa446045b..84fa0e82e7 100644 --- a/internal/frontend/usage/environment.go +++ b/internal/frontend/usage/environment.go @@ -18,11 +18,20 @@ func BudgetEnvironment() []Item { // JobsEnvironment describes the setting the binaries that answer analysis questions // read for how many runs of one plan go concurrently. func JobsEnvironment() []Item { - return []Item{ - {"OPENSYSML_JOBS", "Runs of one check that may go concurrently — the linearizations of an exploration, the engines -engine all consults — each on a worker of its own over the shared model; -jobs and %jobs override it. Default one per CPU, fewer where the memory available leaves less than 512 MiB per worker."}, - } + return []Item{{"OPENSYSML_JOBS", jobsRuns + "; -jobs and %jobs override it. " + jobsDefault}} } +// LoadJobsEnvironment is JobsEnvironment for a binary that also loads files: the +// same count bounds how many files of one load are parsed and validated at once. +func LoadJobsEnvironment() []Item { + return []Item{{"OPENSYSML_JOBS", jobsRuns + ", and files of one load that are parsed and validated at once; -jobs and %jobs override it. " + jobsDefault}} +} + +const ( + jobsRuns = "Runs of one check that may go concurrently — the linearizations of an exploration, the engines -engine all consults — each on a worker of its own over the shared model" + jobsDefault = "Default one per CPU, fewer where the memory available leaves less than 512 MiB per worker." +) + // LegacyPrefixNote states how the superseded variable names are still read, and // belongs with any list of them. const LegacyPrefixNote = "Each variable above also answers to its legacy " + diff --git a/internal/semantic/resolve/document.go b/internal/semantic/resolve/document.go index cc99299f1e..0d2bf7c88c 100644 --- a/internal/semantic/resolve/document.go +++ b/internal/semantic/resolve/document.go @@ -549,22 +549,59 @@ func (r *Resolver) resolveMetadataPrefix(names, parent *symbols.Scope, prefix *a for _, a := range prefix.About { r.ResolveQualified(names, a) } + owner := r.metadataBodyOwner(names, prefix) + body := parent.ChildFor(prefix) + if body == nil { + return + } + // Body values resolve against the metadata definition, not the annotated element. + linkMetadataBody(body, owner) + if owner != nil { + r.resolveMetadataBody(body, prefix.Body) + } +} + +// metadataBodyOwner is the metadata definition the body of prefix resolves against, +// its type read in names; nil when it does not resolve or there is no body. +func (r *Resolver) metadataBodyOwner(names *symbols.Scope, prefix *ast.PrefixMetadata) *symbols.Symbol { owner, ok := r.ResolveQualified(names, prefix.Type) if !ok || owner == nil || len(prefix.Body) == 0 { - return + return nil } if target, aliasOK := r.ResolveAliasTarget(owner); aliasOK { - owner = target + return target } - body := parent.ChildFor(prefix) - if body == nil { + return owner +} + +// linkMetadataBody makes owner, the metadata definition the body resolves against +// now, the body scope's owner; a definition it kept from an earlier build goes. +func linkMetadataBody(body *symbols.Scope, owner *symbols.Symbol) { + if body.Owner() != owner { + body.SetOwner(owner) + } +} + +// LinkMetadataBodies sets every annotation body scope's owner as resolving the +// document would, so resolving it afterwards writes nothing to the scope tree. +func (r *Resolver) LinkMetadataBodies(name string) { + rootScope := r.idx.DocumentRoot(name) + if rootScope == nil { return } - // Body values resolve against the metadata definition, not the annotated element. - if body.Owner() == nil { - body.SetOwner(owner) + saved := r.document + r.document = name + defer func() { r.document = saved }() + r.linkMetadataBodies(rootScope) +} + +func (r *Resolver) linkMetadataBodies(scope *symbols.Scope) { + for _, child := range scope.Children() { + if prefix, ok := child.Node().(*ast.PrefixMetadata); ok { + linkMetadataBody(child, r.metadataBodyOwner(r.bodyScope(scope, child.Annotated()), prefix)) + } + r.linkMetadataBodies(child) } - r.resolveMetadataBody(body, prefix.Body) } func (r *Resolver) resolveMetadataBody(scope *symbols.Scope, members []ast.Node) { diff --git a/internal/semantic/symbols/builder.go b/internal/semantic/symbols/builder.go index fed52541d0..cf2c761154 100644 --- a/internal/semantic/symbols/builder.go +++ b/internal/semantic/symbols/builder.go @@ -54,7 +54,7 @@ func unwrapMember(m ast.Node) (ast.Node, ast.Visibility) { // wrapper before unwrap. func buildDecl(scope *Scope, decl ast.Node, vis ast.Visibility, trivia []ast.Trivia) { if prefixes := prefixMetadataOf(decl); len(prefixes) > 0 { - buildMetadataBodyScopes(scope, prefixes) + buildMetadataBodyScopes(scope, decl, prefixes) } switch { case buildNamespaceDecl(scope, decl, vis, trivia): @@ -177,7 +177,7 @@ func buildBehaviorDecl(scope *Scope, decl ast.Node, vis ast.Visibility, trivia [ // error nodes have no declaration. Nothing to register here. return true case *ast.PrefixMetadata: - child := buildMetadataBodyScope(scope, d) + child := buildMetadataBodyScope(scope, nil, d) // An identification names the usage as a member of its namespace, exactly // as the `metadata` spelling of the same declaration does. if d.Ident.Name != "" || d.Ident.ShortName != "" { @@ -360,21 +360,25 @@ func prefixMetadataOf(decl ast.Node) []*ast.PrefixMetadata { return prefixes } -func buildMetadataBodyScopes(scope *Scope, prefixes []*ast.PrefixMetadata) { +// buildMetadataBodyScopes builds the body scopes of the annotations written on +// decl, a member of scope. +func buildMetadataBodyScopes(scope *Scope, decl ast.Node, prefixes []*ast.PrefixMetadata) { for _, prefix := range prefixes { if prefix != nil && len(prefix.Body) > 0 { - buildMetadataBodyScope(scope, prefix) + buildMetadataBodyScope(scope, decl, prefix) } } } // buildMetadataBodyScope builds the scope of a metadata usage's body and -// returns it, or nil when the usage has no body. -func buildMetadataBodyScope(parent *Scope, prefix *ast.PrefixMetadata) *Scope { +// returns it, or nil when the usage has no body. annotated is the declaration +// the usage is a prefix of, nil for a usage that is a member of its own. +func buildMetadataBodyScope(parent *Scope, annotated ast.Node, prefix *ast.PrefixMetadata) *Scope { if parent == nil || prefix == nil || len(prefix.Body) == 0 { return nil } child := NewScope(parent, prefix) + child.annotated = annotated child.markBodyLocal() parent.AddChild(child) buildMembers(child, prefix.Body) diff --git a/internal/semantic/symbols/index.go b/internal/semantic/symbols/index.go index ff162ecd6f..2a6e2fe04d 100644 --- a/internal/semantic/symbols/index.go +++ b/internal/semantic/symbols/index.go @@ -44,7 +44,7 @@ type Index struct { base *Index frozen bool generation *indexGeneration - directChildrenMu sync.Mutex + directChildrenMu sync.RWMutex directChildrenGeneration uint64 directChildrenCache map[directChildrenKey][]*Symbol directChildrenByName map[directChildrenKey]map[string][]*Symbol @@ -375,7 +375,8 @@ func (idx *Index) AddDocumentWithKind(name string, root *ast.RootNamespace, kind func (idx *Index) addDocument(name string, root *ast.RootNamespace, rs *Scope, kind source.Kind, explicitKind bool) { idx.mustBeWritable("AddDocument") - idx.RemoveDocument(name) + // The caller expands once the documents are in; nothing is read in between. + idx.removeDocument(name, false) idx.changedDoc(name) if rs == nil { rs = Build(root) @@ -826,6 +827,12 @@ func (idx *Index) hasFQN(fqn string, sym *Symbol) bool { // the removal is recorded in the overlay, which stops answering for what the // document contributed while the base keeps it for every other index over it. func (idx *Index) RemoveDocument(name string) { + idx.removeDocument(name, true) +} + +// removeDocument is RemoveDocument, re-expanding only when asked: a replacement +// takes the old document out and expands once the new one is in. +func (idx *Index) removeDocument(name string, expand bool) { idx.mustBeWritable("RemoveDocument") if !idx.knows(name) { return @@ -869,7 +876,9 @@ func (idx *Index) RemoveDocument(name string) { idx.docReexports.del(name) idx.dropNamespaceFilters(name) - idx.ExpandWildcardImports() + if expand { + idx.ExpandWildcardImports() + } } // MarkLibrary records that the named document holds bundled library content, @@ -1600,13 +1609,12 @@ func (idx *Index) ShortNamed(name string) bool { } idx.readSegment(name) generation := idx.generation.get() - idx.directChildrenMu.Lock() - idx.resetDirectChildrenCachesLocked(generation) - if v, ok := idx.shortNamedCache[name]; ok { - idx.directChildrenMu.Unlock() + if v, ok := cachedAt(idx, generation, func() (bool, bool) { + v, ok := idx.shortNamedCache[name] + return v, ok + }); ok { return v } - idx.directChildrenMu.Unlock() v := idx.shortNamedScan(name) idx.directChildrenMu.Lock() if idx.generation.get() == generation { @@ -1704,13 +1712,12 @@ func (idx *Index) LookupDirectChildren(prefix string) []*Symbol { func (idx *Index) lookupDirectChildren(key directChildrenKey) []*Symbol { idx.readNamespace(key.prefix) generation := idx.generation.get() - idx.directChildrenMu.Lock() - idx.resetDirectChildrenCachesLocked(generation) - if out, ok := idx.directChildrenCache[key]; ok { - idx.directChildrenMu.Unlock() + if out, ok := cachedAt(idx, generation, func() ([]*Symbol, bool) { + out, ok := idx.directChildrenCache[key] + return out, ok + }); ok { return out } - idx.directChildrenMu.Unlock() var out []*Symbol keys := idx.childKeys(key.prefix) @@ -1733,6 +1740,23 @@ func (idx *Index) lookupDirectChildren(key directChildrenKey) []*Symbol { return out } +// cachedAt reads a direct-children cache through get under the read lock while +// current for generation; a stale or empty cache takes the write lock and resets. +func cachedAt[V any](idx *Index, generation uint64, get func() (V, bool)) (V, bool) { + idx.directChildrenMu.RLock() + current := idx.directChildrenGeneration == generation + v, ok := get() + idx.directChildrenMu.RUnlock() + if current && ok { + return v, true + } + idx.directChildrenMu.Lock() + idx.resetDirectChildrenCachesLocked(generation) + v, ok = get() + idx.directChildrenMu.Unlock() + return v, ok +} + // resetDirectChildrenCachesLocked drops both direct-children caches when the // index has changed since they were filled. Callers hold directChildrenMu. func (idx *Index) resetDirectChildrenCachesLocked(generation uint64) { @@ -1771,10 +1795,10 @@ func (idx *Index) LookupDirectChildrenNamedFrom(prefix, fromFQN, name string) [] func (idx *Index) lookupDirectChildrenNamed(key directChildrenKey, name string) []*Symbol { idx.readNamespace(key.prefix) generation := idx.generation.get() - idx.directChildrenMu.Lock() - idx.resetDirectChildrenCachesLocked(generation) - byName, ok := idx.directChildrenByName[key] - idx.directChildrenMu.Unlock() + byName, ok := cachedAt(idx, generation, func() (map[string][]*Symbol, bool) { + byName, ok := idx.directChildrenByName[key] + return byName, ok + }) if ok { return byName[name] } diff --git a/internal/semantic/symbols/index_removal_test.go b/internal/semantic/symbols/index_removal_test.go index 0212430674..ef516749f6 100644 --- a/internal/semantic/symbols/index_removal_test.go +++ b/internal/semantic/symbols/index_removal_test.go @@ -281,6 +281,38 @@ func TestEditingTheTargetOfAFileLevelImportDropsItsReexport(t *testing.T) { } } +// Replacing documents that are already indexed expands nothing until the caller +// asks, and that one expansion leaves what a fresh build over the new set gives: +// a reload pays for one expansion, as a first load does. +func TestReplacingDocumentsExpandsOnceEqualToFreshBuild(t *testing.T) { + before := map[string]string{ + "a.sysml": "package Mid { public import Lib::*; part def OldOnly; }", + "b.sysml": "package Top { public import Mid::*; } package Far { public import Top::*; }", + "c.sysml": "package Src { part def Exported; } package User { public import Src::*; }", + } + after := map[string]string{ + "a.sysml": "package Mid { public import Lib::*; part def NewOnly; }", + "b.sysml": "package Top { public import Mid::*; }", + "c.sysml": "package Src { part def Renamed; } package User { public import Src::*; }", + } + reused := buildIndex(t, before) + addDoc(t, reused, "a.sysml", after["a.sysml"]) + // Top's re-export of OldOnly goes at the expansion, not at the replacement. + if len(reused.LookupQualified("Top::OldOnly")) == 0 { + t.Fatal("replacing a.sysml expanded on its own") + } + addDoc(t, reused, "b.sysml", after["b.sysml"]) + addDoc(t, reused, "c.sysml", after["c.sysml"]) + reused.ExpandWildcardImports() + if got := len(reused.LookupQualified("Top::OldOnly")); got != 0 { + t.Errorf("Top::OldOnly = %d symbols after its declaration was replaced, want 0", got) + } + if got, want := indexState(reused), indexState(buildIndex(t, after)); got != want { + t.Errorf("replacing every document left an index a fresh build would not produce:\n%s", + diffLines(want, got)) + } +} + // Deriving every importer from an empty re-export state — what expansion falls // back to when its incremental rounds do not settle — has to rebuild exactly what // was there, including a target that only resolves once another importer has been diff --git a/internal/semantic/symbols/scope.go b/internal/semantic/symbols/scope.go index 2296adeac0..e174a476b4 100644 --- a/internal/semantic/symbols/scope.go +++ b/internal/semantic/symbols/scope.go @@ -22,6 +22,7 @@ type Scope struct { children []*Scope childIndex atomic.Pointer[map[ast.Node]*Scope] // lazily built node -> child scope index for larger scopes bodyLocal bool // declarations live only inside the owning body + annotated ast.Node // for a metadata body, the declaration the annotation is written on docName string // document this scope tree belongs to (stamped by SetDocName) } @@ -68,6 +69,10 @@ func (s *Scope) BodyLocal() bool { return s.bodyLocal } // markBodyLocal records that this scope's names do not escape its body. func (s *Scope) markBodyLocal() { s.bodyLocal = true } +// Annotated returns, for the body scope of a prefix metadata annotation, the +// declaration the annotation is written on; nil for any other scope. +func (s *Scope) Annotated() ast.Node { return s.annotated } + // Children returns the child scopes in definition order. func (s *Scope) Children() []*Scope { return s.children } diff --git a/internal/workspace/model/batch.go b/internal/workspace/model/batch.go new file mode 100644 index 0000000000..089010ea35 --- /dev/null +++ b/internal/workspace/model/batch.go @@ -0,0 +1,164 @@ +package model + +import ( + "bytes" + "errors" + "fmt" + "runtime" + "sync" + "sync/atomic" + + "github.com/Open-MBEE/OpenSysML/internal/check/passes" + "github.com/Open-MBEE/OpenSysML/internal/syntax/diag" +) + +// Input is one document a batch opens: the name it is indexed under, its text +// and its version. +type Input struct { + Name string + Content []byte + Version int +} + +// DefaultWorkers is the worker count a workspace starts with: one per CPU the +// process may run on. +func DefaultWorkers() int { return runtime.GOMAXPROCS(0) } + +// ErrWorkers is the error for a worker count below one. +var ErrWorkers = errors.New("workspace: the worker count must be a positive integer") + +// Workers reports how many documents a batch parses and analyzes at once. +func (w *Workspace) Workers() int { + w.mu.RLock() + defer w.mu.RUnlock() + return w.workers +} + +// SetWorkers sets how many documents a batch parses and analyzes at once. The +// result of a batch is the same at any count; a count below one is an error. +func (w *Workspace) SetWorkers(n int) error { + if n < 1 { + return fmt.Errorf("%w: %d", ErrWorkers, n) + } + w.mu.Lock() + defer w.mu.Unlock() + w.workers = n + return nil +} + +// OpenAll opens the inputs as one batch: parsed on the workers, added to the index +// in order, wildcard imports expanded once. Same result as opening them one by one +// as the batch starts: a document changed by another caller meanwhile keeps that change. +func (w *Workspace) OpenAll(inputs []Input) { + was := w.reserveBatch(inputs) + docs := make([]*Document, len(inputs)) + ParallelFor(w.Workers(), len(inputs), func(i int) { + in := inputs[i] + docs[i] = newDocument(in.Name, bytes.Clone(in.Content), in.Version) + }) + w.commitBatch(was, docs) +} + +// reserveBatch is each input's name's change count as the batch starts, which is +// what commitBatch installs over: a name opened and removed meanwhile is absent +// again, but its count has moved. +func (w *Workspace) reserveBatch(inputs []Input) map[string]uint64 { + w.mu.RLock() + defer w.mu.RUnlock() + was := make(map[string]uint64, len(inputs)) + for _, in := range inputs { + was[in.Name] = w.changes[in.Name] + } + return was +} + +// commitBatch installs the parsed documents whose name is as the batch reserved +// it; a name changed since keeps its newer state. +func (w *Workspace) commitBatch(was map[string]uint64, docs []*Document) { + w.mu.Lock() + defer w.mu.Unlock() + var installed []string + for _, doc := range docs { + if w.changes[doc.Name] != was[doc.Name] { + continue + } + w.open[doc.Name] = true + w.docs[doc.Name] = doc + w.changes[doc.Name]++ + was[doc.Name] = w.changes[doc.Name] + w.installLocked(doc) + installed = append(installed, doc.Name) + } + if len(installed) > 0 { + w.index.ExpandWildcardImports() + w.invalidateLocked(installed...) + } +} + +// DiagnosticsAll returns the named documents' diagnostics in the order named (nil +// for an unknown name), analyzing the uncached ones on the workers, then caching. +// The workers share one gather of the workspace-wide audits, made on first use. +// Pending regathers settle first, so no cached entry they would drop is served. +func (w *Workspace) DiagnosticsAll(names []string) [][]diag.Diagnostic { + w.mu.Lock() + defer w.mu.Unlock() + w.settleGathersLocked() + out := make([][]diag.Diagnostic, len(names)) + var pending []string + queued := map[string]bool{} + for i, name := range names { + if w.docs[name] == nil { + continue + } + if cached, ok := w.diagCache[name]; ok { + out[i] = cached + } else if !queued[name] { + queued[name] = true + pending = append(pending, name) + } + } + batch := &passes.Batch{Documents: pending, Gathers: passes.NewGathers(), Source: w.sourceText()} + passes.PrepareBatch(w.index, batch) + analyzed := make([][]diag.Diagnostic, len(pending)) + ParallelFor(w.workers, len(pending), func(i int) { + analyzed[i] = w.analyze(pending[i], w.docs[pending[i]], batch) + }) + for i, name := range pending { + w.diagCache[name] = analyzed[i] + w.batched[name] = true + } + for i, name := range names { + if out[i] == nil && w.docs[name] != nil { + out[i] = w.diagCache[name] + } + } + return out +} + +// ParallelFor runs fn(i) for every i below n on up to workers goroutines, and +// returns once every call has; it is how a batch spreads its documents. +func ParallelFor(workers, n int, fn func(i int)) { + workers = min(workers, n) + if workers <= 1 { + for i := 0; i < n; i++ { + fn(i) + } + return + } + var next atomic.Int64 + var wg sync.WaitGroup + for range workers { + wg.Add(1) + go func() { + defer wg.Done() + for { + i := int(next.Add(1) - 1) + if i >= n { + return + } + fn(i) + } + }() + } + wg.Wait() +} diff --git a/internal/workspace/model/batch_test.go b/internal/workspace/model/batch_test.go new file mode 100644 index 0000000000..122bb6c188 --- /dev/null +++ b/internal/workspace/model/batch_test.go @@ -0,0 +1,458 @@ +package model + +import ( + "errors" + "fmt" + "os" + "path/filepath" + "sort" + "strings" + "testing" + + "github.com/Open-MBEE/OpenSysML/internal/syntax/diag" + "github.com/Open-MBEE/OpenSysML/internal/syntax/source" +) + +// The model roots a batch is checked against: the fixtures, the shipped examples +// and the four OMG corpora, the latter skipped when absent unless CI requires them. +var batchRoots = []struct { + dir string + require string + fetch string +}{ + {dir: "../../../tests/testdata"}, + {dir: "../../../examples"}, + {dir: "../../../examples/sysml-v2-training", require: "OPENSYSML_REQUIRE_TRAINING_CORPUS", fetch: "./scripts/download-training-examples.sh"}, + {dir: "../../../examples/pilot-corpora/kerml-examples", require: "OPENSYSML_REQUIRE_PILOT_CORPORA", fetch: "./scripts/download-pilot-corpora.sh"}, + {dir: "../../../examples/pilot-corpora/sysml-examples", require: "OPENSYSML_REQUIRE_PILOT_CORPORA", fetch: "./scripts/download-pilot-corpora.sh"}, + {dir: "../../../examples/pilot-corpora/sysml-validation", require: "OPENSYSML_REQUIRE_PILOT_CORPORA", fetch: "./scripts/download-pilot-corpora.sh"}, +} + +// Every directory of models, opened as one batch, reports the same diagnostics in +// the same order on one worker as on many, and as opening its files one by one. +func TestParallelBatchValidationMatchesSerial(t *testing.T) { + seen := map[string]bool{} + for _, root := range batchRoots { + if _, err := os.Stat(root.dir); os.IsNotExist(err) { + if os.Getenv(root.require) != "" { + t.Fatalf("%s is missing and %s is set; fetch it with %s", root.dir, root.require, root.fetch) + } + t.Logf("%s is absent (fetch it with %s), so this run proves nothing about it", root.dir, root.fetch) + continue + } + for _, files := range modelDirectories(t, root.dir) { + dir := filepath.Dir(files[0]) + if seen[dir] { + continue + } + seen[dir] = true + t.Run(filepath.ToSlash(dir), func(t *testing.T) { + inputs := readInputs(t, files) + serial := serialDiagnostics(inputs) + for _, workers := range []int{1, 2, 8} { + if got := batchDiagnostics(t, inputs, workers); got != serial { + t.Errorf("a batch on %d workers reported:\n%s\nwant, as opening the files one by one does:\n%s", workers, got, serial) + } + } + }) + } + } +} + +// A batch analyzes only what is not cached, answers every name asked for in the +// order asked, repeats included, and nil for a name the workspace does not hold. +func TestDiagnosticsAllAnswersInTheOrderAsked(t *testing.T) { + ws := NewWorkspace() + ws.OpenAll([]Input{ + {Name: "a.sysml", Content: []byte("package A { part def X; }"), Version: 1}, + {Name: "b.sysml", Content: []byte("package B { part x : A::X; part y : Missing; }"), Version: 1}, + }) + if diags := ws.Diagnostics("a.sysml"); len(diags) != 0 { + t.Fatalf("a.sysml should analyze cleanly, got %+v", diags) + } + got := ws.DiagnosticsAll([]string{"b.sysml", "none.sysml", "a.sysml", "b.sysml"}) + if len(got) != 4 || got[1] != nil || len(got[2]) != 0 { + t.Fatalf("DiagnosticsAll = %v, want b, nil, none, b", got) + } + if len(got[0]) != 1 || got[0][0].Message != "unresolved reference: Missing" || len(got[3]) != 1 { + t.Fatalf("b.sysml reported %v, want one unresolved reference to Missing", got[0]) + } + if &got[0][0] != &got[3][0] { + t.Errorf("the two answers for b.sysml should be the one cached slice") + } +} + +// A batch whose annotation types are reached through inherited and filtered +// imports, typed usages and another document — the lookups only a model-backed +// resolver answers — reports on many workers what one reports; under the race +// detector this also shows the workers writing nothing into the shared scopes. +func TestParallelBatchLinksAnnotationsFoundThroughTheModel(t *testing.T) { + inputs := []Input{ + {Name: "meta.sysml", Version: 1, Content: []byte(`package Meta { + metadata def Tag { attribute n; } + metadata def Other { attribute m; } +}`)}, + {Name: "model.sysml", Version: 1, Content: []byte(`package P { + part def Base { public import Meta::*; } + part def Sub :> Base; + part def Filtered { public import Meta::*[@Meta::Tag]; } + part b : Base; + part def Owned { metadata def Own { attribute n; } } + part def Derived :> Owned; + part def C { + @Meta::Tag { n = 1; } + @Sub::Tag { n = 2; } + @Filtered::Tag { n = 3; } + @b::Tag { n = 4; } + @Derived::Own { n = 5; } + } +}`)}, + {Name: "audit.sysml", Version: 1, Content: []byte(`package Audit { + part def Ground { part c : P::C; } + part sat : Ground; +}`)}, + } + serial := serialDiagnostics(inputs) + for range 8 { + if got := batchDiagnostics(t, inputs, 8); got != serial { + t.Fatalf("a batch on 8 workers reported:\n%s\nwant, as opening the files one by one does:\n%s", got, serial) + } + } +} + +// Asking for some of a batch's documents leaves the others' scopes as they are: +// the workspace-wide passes read them, so the answer is what asking one by one gives. +func TestDiagnosticsAllOverSomeDocumentsMatchesAskingOneByOne(t *testing.T) { + inputs := []Input{ + {Name: "meta.sysml", Version: 1, Content: []byte(`package Meta { + metadata def Tag { attribute n; } + part def Base { public import Meta::*; } + part def Sub :> Base; + part def C { + @Meta::Tag { n = 1; } + @Sub::Tag { n = 2; } + } + part def D; + metadata Tag about D { n = 3; } +}`)}, + {Name: "a.sysml", Version: 1, Content: []byte(`package A { + part def Ground { part c : Meta::C; @Meta::Tag { n = 4; } } +}`)}, + {Name: "b.sysml", Version: 1, Content: []byte(`package B { + part def Station { part c : Meta::C; @Meta::Tag { n = 5; } } +}`)}, + } + asked := []string{"a.sysml", "b.sysml"} + serial := NewWorkspace() + for _, in := range inputs { + serial.Open(in.Name, in.Content, in.Version) + } + var want strings.Builder + for _, name := range asked { + renderDiagnostics(&want, name, serial.Diagnostics(name)) + } + for range 8 { + ws := NewWorkspace() + if err := ws.SetWorkers(8); err != nil { + t.Fatal(err) + } + ws.OpenAll(inputs) + var got strings.Builder + for i, diags := range ws.DiagnosticsAll(asked) { + renderDiagnostics(&got, asked[i], diags) + } + if got.String() != want.String() { + t.Fatalf("asking for two of three documents on 8 workers reported:\n%s\nwant, as asking one by one does:\n%s", got.String(), want.String()) + } + } +} + +// A batch of one document prepares the others' annotation bodies too: the +// identity gather files an unlinked body's attribute under its bare name, where +// it would collide with a declared id of the document asked for. +func TestDiagnosticsAllOverOneDocumentPreparesTheOthersBodies(t *testing.T) { + meta := Input{Name: "meta.sysml", Version: 1, Content: []byte(`package Meta { + metadata def M { attribute rationale : ScalarValues::String; attribute fallback : ScalarValues::String = "d"; } + part def X { @M { rationale = fallback; } } +}`)} + check := Input{Name: "check.sysml", Version: 1, Content: []byte(`package Check { + part def Y { @IdentityMetadata::ElementId { id = "rationale"; } } +}`)} + inputs := []Input{meta, check} + serial := NewWorkspace() + for _, in := range inputs { + serial.Open(in.Name, in.Content, in.Version) + } + serial.Diagnostics(meta.Name) + var want strings.Builder + renderDiagnostics(&want, check.Name, serial.Diagnostics(check.Name)) + if !strings.Contains(want.String(), "identity-unscoped-id") || strings.Contains(want.String(), "identity-duplicate-id") { + t.Fatalf("check.sysml over resolved documents should report only the unscoped id, got:\n%s", want.String()) + } + for _, workers := range []int{1, 8} { + ws := NewWorkspace() + if err := ws.SetWorkers(workers); err != nil { + t.Fatal(err) + } + ws.OpenAll(inputs) + var got strings.Builder + renderDiagnostics(&got, check.Name, ws.DiagnosticsAll([]string{check.Name})[0]) + if got.String() != want.String() { + t.Errorf("asking for check.sysml alone on %d workers reported:\n%s\nwant:\n%s", workers, got.String(), want.String()) + } + } +} + +// A batch's contexts read comment bodies from the documents as the editor's +// model does, so an import filtered on Comment::body hides the same names on +// both paths instead of keeping every candidate as an unevaluable filter would. +func TestDiagnosticsAllReadsCommentBodiesAsAnEditorDoes(t *testing.T) { + inputs := []Input{ + {Name: "p.sysml", Version: 1, Content: []byte(`package P { + comment Shown /* public */ + comment Hidden /* private */ +}`)}, + {Name: "a.sysml", Version: 1, Content: []byte(`package A { + private import KerML::*; + private import P::*[Comment::body == "public"]; + comment about Shown /* seen */ + comment about Hidden /* filtered out */ + private import Q::Hidden; +} +package Q { + private import KerML::*; + public import P::*; + filter Comment::body == "public"; +}`)}, + } + serial := NewWorkspace() + for _, in := range inputs { + serial.Open(in.Name, in.Content, in.Version) + } + var want strings.Builder + renderDiagnostics(&want, "a.sysml", serial.Diagnostics("a.sysml")) + if n := strings.Count(want.String(), "unresolved reference"); n != 2 { + t.Fatalf("an editor should report Hidden and Q::Hidden unresolved, got:\n%s", want.String()) + } + for _, workers := range []int{1, 8} { + ws := NewWorkspace() + if err := ws.SetWorkers(workers); err != nil { + t.Fatal(err) + } + ws.OpenAll(inputs) + var got strings.Builder + renderDiagnostics(&got, "a.sysml", ws.DiagnosticsAll([]string{"a.sysml"})[0]) + if got.String() != want.String() { + t.Errorf("a batch on %d workers reported:\n%s\nwant, as the editor does:\n%s", workers, got.String(), want.String()) + } + } +} + +// A batch opened over documents already there replaces them as Open does, so +// the index holds each name once. +func TestOpenAllReplacesEarlierDocuments(t *testing.T) { + ws := NewWorkspace() + ws.Open("a.sysml", []byte("package A { part def Old; }"), 1) + ws.OpenAll([]Input{{Name: "a.sysml", Content: []byte("package A { part def New; }"), Version: 2}}) + if syms := ws.LookupQualified("A::Old"); len(syms) != 0 { + t.Errorf("A::Old should be gone, found %d", len(syms)) + } + if syms := ws.LookupQualified("A::New"); len(syms) != 1 { + t.Errorf("A::New should be indexed once, found %d", len(syms)) + } + if doc := ws.Document("a.sysml"); doc == nil || doc.Version != 2 || !ws.IsOpen("a.sysml") { + t.Errorf("a.sysml should be the open version 2 document, got %+v", doc) + } +} + +// A document another caller changes while a batch parses keeps that change: the +// batch installs only over what it reserved, so an edit, a buffer opened, a +// removal and an open-then-remove made meanwhile all stand, and only the +// untouched name is opened. +func TestOpenAllKeepsAChangeMadeWhileItParsed(t *testing.T) { + ws := NewWorkspace() + ws.Open("a.sysml", []byte("package A { part def Old; }"), 1) + ws.Open("d.sysml", []byte("package D { part def Old; }"), 1) + inputs := []Input{ + {Name: "a.sysml", Content: []byte("package A { part def Batch; }"), Version: 2}, + {Name: "b.sysml", Content: []byte("package B { part def Batch; }"), Version: 1}, + {Name: "c.sysml", Content: []byte("package C { part def Batch; }"), Version: 1}, + {Name: "d.sysml", Content: []byte("package D { part def Batch; }"), Version: 2}, + {Name: "e.sysml", Content: []byte("package E { part def Batch; }"), Version: 1}, + } + was := ws.reserveBatch(inputs) + docs := make([]*Document, len(inputs)) + for i, in := range inputs { + docs[i] = newDocument(in.Name, in.Content, in.Version) + } + ws.Update("a.sysml", []byte("package A { part def Edited; }"), 3) + ws.Open("b.sysml", []byte("package B { part def Opened; }"), 1) + ws.Remove("d.sysml") + ws.Open("e.sysml", []byte("package E { part def Opened; }"), 1) + ws.Remove("e.sysml") + ws.commitBatch(was, docs) + + if doc := ws.Document("a.sysml"); doc == nil || doc.Version != 3 { + t.Errorf("a.sysml should keep the edit made while the batch parsed, got %+v", doc) + } + if syms := ws.LookupQualified("A::Edited"); len(syms) != 1 { + t.Errorf("A::Edited should be indexed once, found %d", len(syms)) + } + if syms := ws.LookupQualified("A::Batch"); len(syms) != 0 { + t.Errorf("the batch's stale a.sysml should not be indexed, found A::Batch %d times", len(syms)) + } + if syms := ws.LookupQualified("B::Opened"); len(syms) != 1 || len(ws.LookupQualified("B::Batch")) != 0 { + t.Error("b.sysml was opened while the batch parsed and should keep that buffer") + } + if ws.Document("d.sysml") != nil || len(ws.LookupQualified("D::Batch")) != 0 { + t.Error("d.sysml was removed while the batch parsed and should stay removed") + } + if ws.Document("e.sysml") != nil || len(ws.LookupQualified("E::Batch")) != 0 { + t.Error("e.sysml was opened and removed while the batch parsed and should stay removed") + } + if doc := ws.Document("c.sysml"); doc == nil || !ws.IsOpen("c.sysml") { + t.Errorf("c.sysml, untouched meanwhile, should be opened by the batch, got %+v", doc) + } + if syms := ws.LookupQualified("C::Batch"); len(syms) != 1 { + t.Errorf("C::Batch should be indexed once, found %d", len(syms)) + } +} + +func TestWorkersSetting(t *testing.T) { + ws := NewWorkspace() + if ws.Workers() != DefaultWorkers() || DefaultWorkers() < 1 { + t.Fatalf("Workers() = %d, want the default %d", ws.Workers(), DefaultWorkers()) + } + if err := ws.SetWorkers(0); !errors.Is(err, ErrWorkers) { + t.Errorf("SetWorkers(0) = %v, want ErrWorkers", err) + } + if err := ws.SetWorkers(3); err != nil || ws.Workers() != 3 { + t.Errorf("SetWorkers(3) = %v, Workers() = %d", err, ws.Workers()) + } +} + +// modelDirectories walks root and returns the model files of every directory +// holding one or more, each directory's files sorted. +func modelDirectories(t *testing.T, root string) [][]string { + t.Helper() + byDir := map[string][]string{} + err := filepath.WalkDir(root, func(path string, entry os.DirEntry, err error) error { + if err != nil { + return err + } + if entry.IsDir() || source.KindOf(path) == source.KindUnknown { + return nil + } + byDir[filepath.Dir(path)] = append(byDir[filepath.Dir(path)], path) + return nil + }) + if err != nil { + t.Fatalf("scan %s: %v", root, err) + } + dirs := make([]string, 0, len(byDir)) + for dir := range byDir { + dirs = append(dirs, dir) + } + sort.Strings(dirs) + out := make([][]string, 0, len(dirs)) + for _, dir := range dirs { + files := byDir[dir] + sort.Strings(files) + out = append(out, files) + } + return out +} + +func readInputs(t *testing.T, paths []string) []Input { + t.Helper() + inputs := make([]Input, 0, len(paths)) + for _, path := range paths { + content, err := os.ReadFile(path) + if err != nil { + t.Fatal(err) + } + inputs = append(inputs, Input{Name: path, Content: content, Version: 1}) + } + return inputs +} + +// serialDiagnostics opens the inputs one by one and diagnoses each in turn, the +// path the editor takes, and renders the result for comparison. +func serialDiagnostics(inputs []Input) string { + ws := NewWorkspace() + for _, in := range inputs { + ws.Open(in.Name, in.Content, in.Version) + } + var b strings.Builder + for _, in := range inputs { + renderDiagnostics(&b, in.Name, ws.Diagnostics(in.Name)) + } + return b.String() +} + +// batchDiagnostics opens the inputs as one batch on the given workers and +// diagnoses them as one batch, rendering the result as serialDiagnostics does. +func batchDiagnostics(t *testing.T, inputs []Input, workers int) string { + t.Helper() + ws := NewWorkspace() + if err := ws.SetWorkers(workers); err != nil { + t.Fatal(err) + } + ws.OpenAll(inputs) + names := make([]string, len(inputs)) + for i, in := range inputs { + names[i] = in.Name + } + var b strings.Builder + for i, diags := range ws.DiagnosticsAll(names) { + renderDiagnostics(&b, names[i], diags) + } + return b.String() +} + +// renderDiagnostics writes every field of each diagnostic, in the order given, +// so a comparison sees content and order alike. +func renderDiagnostics(b *strings.Builder, name string, diags []diag.Diagnostic) { + for _, d := range diags { + fmt.Fprintf(b, "%s: %+v\n", filepath.Base(name), d) + } +} + +// A batch that reloads a metadata definition's document leaves the annotating +// documents in place; the next batch reports them over the definition now +// indexed, on any number of workers, as a fresh workspace does. +func TestOpenAllReloadedMetadataDefinitionReownsAnnotationBodies(t *testing.T) { + use := Input{Name: "use.sysml", Version: 1, Content: []byte("package Use { private import Meta::*; part def C { @M { a = b; } } }")} + first := Input{Name: "meta.sysml", Version: 1, Content: []byte("package Meta { metadata def M { attribute a; attribute b; } }")} + edited := Input{Name: "meta.sysml", Version: 2, Content: []byte("package Meta { metadata def M { attribute a; attribute c; } }")} + names := []string{"meta.sysml", "use.sysml"} + fresh := NewWorkspace() + fresh.OpenAll([]Input{edited, use}) + var want strings.Builder + for i, diags := range fresh.DiagnosticsAll(names) { + renderDiagnostics(&want, names[i], diags) + } + if !strings.Contains(want.String(), "unresolved reference: b") { + t.Fatalf("a fresh workspace should report b, which M no longer declares, got:\n%s", want.String()) + } + for _, workers := range []int{1, 8} { + ws := NewWorkspace() + if err := ws.SetWorkers(workers); err != nil { + t.Fatal(err) + } + ws.OpenAll([]Input{first, use}) + for i, diags := range ws.DiagnosticsAll(names) { + if len(diags) != 0 { + t.Fatalf("%d workers, before the reload, %s: %v", workers, names[i], diags) + } + } + ws.OpenAll([]Input{edited}) + var got strings.Builder + for i, diags := range ws.DiagnosticsAll(names) { + renderDiagnostics(&got, names[i], diags) + } + if got.String() != want.String() { + t.Errorf("%d workers, after reloading M:\n%s\nwant, as a fresh workspace reports:\n%s", workers, got.String(), want.String()) + } + } +} diff --git a/internal/workspace/model/lazy_regather_test.go b/internal/workspace/model/lazy_regather_test.go index 9b727947a9..faf2bb6aa1 100644 --- a/internal/workspace/model/lazy_regather_test.go +++ b/internal/workspace/model/lazy_regather_test.go @@ -38,3 +38,27 @@ func TestWorkspacePendingGathersSettleOnRead(t *testing.T) { t.Errorf("sat.sysml: incremental %v, fresh %v", gotSat, want) } } + +// TestBatchSettlesPendingGathersBeforeServingTheCache: a batch consulted after +// an edit settles the regathers the edit queued before it serves a cached +// verdict, so the entry the settle would drop is not the answer. +func TestBatchSettlesPendingGathersBeforeServingTheCache(t *testing.T) { + ws := NewWorkspace() + ws.Open("hub.sysml", []byte("package M { private import OOSEM::*; #stakeholderNeed requirement need; }"), 1) + ws.Open("sat.sysml", []byte("package S { private import OOSEM::*; #systemRequirement requirement sys; }"), 1) + if got := codesOf(ws.Diagnostics("sat.sysml")); len(got) != 0 { + t.Fatalf("sat.sysml under a stakeholder need: %v, want clean", got) + } + + // The hub's need becomes a mission requirement with no read in between; the + // satellite's cached verdict stands until the pending regather runs. + ws.Update("hub.sysml", []byte("package M { private import OOSEM::*; #missionRequirement requirement mission; }"), 2) + got := codesOf(ws.DiagnosticsAll([]string{"sat.sysml"})[0]) + + fresh := NewWorkspace() + fresh.Open("hub.sysml", ws.Document("hub.sysml").Content, 1) + fresh.Open("sat.sysml", ws.Document("sat.sysml").Content, 1) + if want := codesOf(fresh.Diagnostics("sat.sysml")); !reflect.DeepEqual(got, want) { + t.Errorf("sat.sysml: batch after the edit %v, fresh %v", got, want) + } +} diff --git a/internal/workspace/model/workspace.go b/internal/workspace/model/workspace.go index 467dabb9b0..3da5afe079 100644 --- a/internal/workspace/model/workspace.go +++ b/internal/workspace/model/workspace.go @@ -21,11 +21,14 @@ import ( // document set plus the global symbol index. Mutations are serialized under a // write lock; reads take a read lock. type Workspace struct { - mu sync.RWMutex - docs map[string]*Document - onDisk map[string][]byte // last-known on-disk bytes, used when a doc is not open - open map[string]bool // names with an authoritative open buffer - index *symbols.Index + mu sync.RWMutex + docs map[string]*Document + // changes counts the times each name's document was installed or removed, + // so a batch can tell a name changed under it even when it is absent again. + changes map[string]uint64 + onDisk map[string][]byte // last-known on-disk bytes, used when a doc is not open + open map[string]bool // names with an authoritative open buffer + index *symbols.Index // library is every library file the index held at construction, so a displaced // one can come back; libraryRoots names their top-level packages. library map[string]libraryFile @@ -66,6 +69,12 @@ type Workspace struct { // displaced, by standing in for it or by taking its name, for when it comes back. standIns map[string]string displaced map[string]symbols.LibraryDocument + + // workers is how many documents OpenAll and DiagnosticsAll work on at once. + workers int + // batched names the diagCache entries a batch computed: in contexts of their + // own, recording no dependencies, so any change drops them all. + batched map[string]bool } // libraryFile is a library file as indexed: its parsed root, language and mark. @@ -107,6 +116,7 @@ func NewWorkspace(opts ...Option) *Workspace { func NewWorkspaceWithIndex(idx *symbols.Index, opts ...Option) *Workspace { w := &Workspace{ docs: map[string]*Document{}, + changes: map[string]uint64{}, onDisk: map[string][]byte{}, open: map[string]bool{}, index: idx, @@ -117,6 +127,8 @@ func NewWorkspaceWithIndex(idx *symbols.Index, opts ...Option) *Workspace { libDocs: map[string]*Document{}, standIns: map[string]string{}, displaced: map[string]symbols.LibraryDocument{}, + workers: DefaultWorkers(), + batched: map[string]bool{}, } for _, name := range idx.Documents() { record := idx.LibraryDocumentOf(name) @@ -321,39 +333,52 @@ func (w *Workspace) Remove(name string) { func (w *Workspace) reindexLocked(name string, content []byte, version int) { doc := newDocument(name, content, version) w.docs[name] = doc - w.displaceLocked(name) - w.index.AddDocumentScope(name, doc.AST, doc.Scope) // removes stale entries first - w.standInLocked(name, doc) - w.index.ExpandWildcardImports() // Expand new document's wildcard imports + w.changes[name]++ + w.installLocked(doc) + w.index.ExpandWildcardImports() w.invalidateLocked(name) } +// installLocked indexes doc over its built scope, displacing the bundled library +// file of its name and standing in for the one it is a version of; the caller +// expands wildcard imports once its documents are in. Caller holds the write lock. +func (w *Workspace) installLocked(doc *Document) { + w.displaceLocked(doc.Name) + w.index.AddDocumentScope(doc.Name, doc.AST, doc.Scope) // removes stale entries first + w.standInLocked(doc.Name, doc) +} + // removeLocked drops name from the document set and index. Caller holds the lock. func (w *Workspace) removeLocked(name string) { delete(w.docs, name) + w.changes[name]++ w.index.RemoveDocument(name) w.releaseStandInLocked(name) w.restoreLocked(name) w.invalidateLocked(name) } -// invalidateLocked drops what the replacement of name made stale: the resolver -// and model entries name and its dependents own, transitively, with their -// diagnostics and reverse references. Caller holds the write lock. -func (w *Workspace) invalidateLocked(name string) { +// invalidateLocked drops what the replacement of names made stale: the resolver +// and model entries they and their dependents own, transitively, with their +// diagnostics and reverse references, and every batch-computed diagnostic. +// Caller holds the write lock. +func (w *Workspace) invalidateLocked(names ...string) { if w.resolver == nil { w.invalidateAllLocked() return } + w.dropBatchedLocked() w.generation++ ch := w.index.TakeChanges() if ch.Docs == nil { ch.Docs = map[string]bool{} } - ch.Docs[name] = true + for _, name := range names { + ch.Docs[name] = true + delete(w.diagCache, name) + w.refs.drop(name) + } dropped := w.resolver.Invalidate(ch) - delete(w.diagCache, name) - w.refs.drop(name) // The gathers the drop took are replayed by the next read that needs them // (settleGathersLocked), not inside the edit. if w.regatherPending == nil { @@ -403,6 +428,14 @@ func (w *Workspace) settleGathersLocked() { } } +// dropBatchedLocked forgets the diagnostics batches computed. Caller holds the write lock. +func (w *Workspace) dropBatchedLocked() { + for name := range w.batched { + delete(w.diagCache, name) + } + w.batched = map[string]bool{} +} + // contextLocked is a pass context over the workspace's shared semantic state, // for work done between analyses. Caller holds the write lock. func (w *Workspace) contextLocked() *passes.Context { @@ -415,6 +448,7 @@ func (w *Workspace) contextLocked() *passes.Context { // all: the conformance mode. Caller holds the write lock. func (w *Workspace) invalidateAllLocked() { w.diagCache = map[string][]diag.Diagnostic{} + w.batched = map[string]bool{} w.refs = nil w.regatherPending = nil w.generation++ @@ -456,6 +490,14 @@ func (w *Workspace) diagnosticsLocked(name string, doc *Document) []diag.Diagnos if cached, ok := w.diagCache[name]; ok { return cached } + diags := w.analyze(name, doc, nil) + w.diagCache[name] = diags + return diags +} + +// analyze runs the passes over doc with a context of its own, reading the index, +// the analysis options and the batch only, so documents can be analyzed at once. +func (w *Workspace) analyze(name string, doc *Document, batch *passes.Batch) []diag.Diagnostic { parseDiags := make([]diag.Diagnostic, 0, len(doc.ParseDiagnostics)+len(doc.ParseWarnings)) for _, pd := range doc.ParseDiagnostics { parseDiags = append(parseDiags, diag.Diagnostic{ @@ -477,9 +519,10 @@ func (w *Workspace) diagnosticsLocked(name string, doc *Document) []diag.Diagnos Fixes: pw.Fixes, }) } - diags := passes.AnalyzeShared(name, source.KindOf(name), doc.AST, parseDiags, w.analysis, w.sharedLocked()) - w.diagCache[name] = diags - return diags + if batch != nil { + return passes.AnalyzeInBatch(name, source.KindOf(name), doc.AST, parseDiags, w.index, w.analysis, batch) + } + return passes.AnalyzeShared(name, source.KindOf(name), doc.AST, parseDiags, w.analysis, w.sharedLocked()) } // LookupQualified resolves a fully-qualified name against the global index under diff --git a/internal/workspace/model/workspace_test.go b/internal/workspace/model/workspace_test.go index 67d92ef104..8c36801389 100644 --- a/internal/workspace/model/workspace_test.go +++ b/internal/workspace/model/workspace_test.go @@ -193,3 +193,28 @@ func codesOf(diags []diag.Diagnostic) map[string]int { } return out } + +// An annotation body resolves its values against the metadata definition it +// names; when the definition's document changes, an unchanged annotating +// document reports over the definition now indexed, as a fresh workspace does. +func TestWorkspaceReloadedMetadataDefinitionReownsAnnotationBodies(t *testing.T) { + use := []byte("package Use { private import Meta::*; part def C { @M { a = b; } } }") + ws := NewWorkspace() + ws.Open("meta.sysml", []byte("package Meta { metadata def M { attribute a; attribute b; } }"), 1) + ws.Open("use.sysml", use, 1) + if diags := ws.Diagnostics("use.sysml"); len(diags) != 0 { + t.Fatalf("before the edit: %v", diags) + } + edited := []byte("package Meta { metadata def M { attribute a; attribute c; } }") + ws.Update("meta.sysml", edited, 2) + fresh := NewWorkspace() + fresh.Open("meta.sysml", edited, 1) + fresh.Open("use.sysml", use, 1) + want := fresh.Diagnostics("use.sysml") + if len(want) == 0 { + t.Fatal("a fresh workspace should report b, which M no longer declares") + } + if got := ws.Diagnostics("use.sysml"); !reflect.DeepEqual(got, want) { + t.Errorf("after editing M: %v\nwant, as a fresh workspace reports: %v", got, want) + } +} diff --git a/packaging/man/man1/sysml.1 b/packaging/man/man1/sysml.1 index 30a5ffe038..84b53e449c 100644 --- a/packaging/man/man1/sysml.1 +++ b/packaging/man/man1/sysml.1 @@ -132,9 +132,10 @@ Draw this many values uniformly from each \-sweep range instead of stepping through it; needs \-seed .TP .BR \-jobs " \fIn\fP" -Runs of one check that may go concurrently, each on a worker of its own; -default OPENSYSML_JOBS, else one per CPU, fewer where the memory available -leaves less than 512 MiB per worker +Runs of one check that may go concurrently, each on a worker of its own, and +files of one load that are parsed and validated at once; the result is the +same at any count. Default OPENSYSML_JOBS, else one per CPU, fewer where the +memory available leaves less than 512 MiB per worker .SS Analysis engines .TP .B \-engines @@ -479,9 +480,10 @@ its connectors. .PP \-jobs is how many runs of one check may go concurrently \(em the linearizations of an exploration, the engines \-engine all consults \(em each -on a worker of its own over the shared model; the result is the same at any -count. By default one per CPU, fewer where the memory available leaves less -than 512 MiB per worker. +on a worker of its own over the shared model, and how many files of one load +are parsed and validated at once; the result is the same at any count. By +default one per CPU, fewer where the memory available leaves less than 512 MiB +per worker. .SH ANALYSIS ENGINES .RS 2 .nf @@ -976,8 +978,9 @@ of its own with its own budgets. Default 1000. .B OPENSYSML_JOBS Runs of one check that may go concurrently \(em the linearizations of an exploration, the engines \-engine all consults \(em each on a worker of its -own over the shared model; \-jobs and %jobs override it. Default one per CPU, -fewer where the memory available leaves less than 512 MiB per worker. +own over the shared model, and files of one load that are parsed and validated +at once; \-jobs and %jobs override it. Default one per CPU, fewer where the +memory available leaves less than 512 MiB per worker. .TP .B OPENSYSML_TOOLS Directory of the tool manifest: one JSON file per external tool (toolName, diff --git a/tests/resolve/link_metadata_test.go b/tests/resolve/link_metadata_test.go new file mode 100644 index 0000000000..8fab652668 --- /dev/null +++ b/tests/resolve/link_metadata_test.go @@ -0,0 +1,232 @@ +package resolve_test + +import ( + "reflect" + "testing" + + "github.com/Open-MBEE/OpenSysML/internal/semantic/resolve" + "github.com/Open-MBEE/OpenSysML/internal/semantic/semantics" + "github.com/Open-MBEE/OpenSysML/internal/semantic/symbols" + "github.com/Open-MBEE/OpenSysML/internal/syntax/ast" + "github.com/Open-MBEE/OpenSysML/internal/syntax/parser" + "github.com/Open-MBEE/OpenSysML/internal/syntax/source" +) + +// linkedMetadataBodies covers how an annotation body finds its owner: in a +// definition, at the root, about an element, via an alias, nested, unresolved. +const linkedMetadataBodies = `package P { + attribute def T { attribute b; } + metadata def M { attribute a : T; attribute n; } + alias N for M; + part def C { + @M { n = 1; a { b = 2; } } + @N { n = 3; } + @Missing { n = 4; } + } + @M about C { n = 5; } + @N { n = 6; } +}` + +// ownersOf renders the owner of every scope of the tree in tree order, "" for none. +func ownersOf(idx *symbols.Index, root *symbols.Scope) []string { + var out []string + var visit func(*symbols.Scope) + visit = func(s *symbols.Scope) { + owner := "" + if s.Owner() != nil { + owner = idx.GetFQN(s.Owner()) + } + out = append(out, owner) + for _, c := range s.Children() { + visit(c) + } + } + visit(root) + return out +} + +func indexedDoc(t *testing.T, name, src string) (*symbols.Index, *ast.RootNamespace, *resolve.Resolver) { + t.Helper() + p := parser.New(source.New(name, []byte(src))) + root := p.ParseFile() + if len(p.Diagnostics) != 0 { + t.Fatalf("parse diagnostics: %v", p.Diagnostics) + } + idx := symbols.NewIndexFromDoc(name, root) + idx.ExpandWildcardImports() + r := resolve.New(idx) + r.SetModel(semantics.NewModel(r)) + return idx, root, r +} + +// Linking a document's metadata bodies sets exactly the owners resolving it sets, +// so resolving a linked document changes nothing and reports what it would have. +func TestLinkMetadataBodiesSetsWhatResolvingWould(t *testing.T) { + const name = "linked.sysml" + for _, src := range []string{linkedMetadataBodies, nestedMetadataBodies["nested.sysml"], nestedMetadataBodies["nested.kerml"]} { + idx, root, r := indexedDoc(t, name, src) + r.ResolveDocument(name, root) + want := ownersOf(idx, idx.DocumentRoot(name)) + + linkedIdx, linkedRoot, linker := indexedDoc(t, name, src) + linker.LinkMetadataBodies(name) + linked := ownersOf(linkedIdx, linkedIdx.DocumentRoot(name)) + if !reflect.DeepEqual(linked, want) { + t.Errorf("linked owners\n%q\nwant those resolving sets\n%q", linked, want) + } + resolver := resolve.New(linkedIdx) + resolver.SetModel(semantics.NewModel(resolver)) + resolver.ResolveDocument(name, linkedRoot) + if got := ownersOf(linkedIdx, linkedIdx.DocumentRoot(name)); !reflect.DeepEqual(got, linked) { + t.Errorf("resolving a linked document changed owners to\n%q\nfrom\n%q", got, linked) + } + if !reflect.DeepEqual(resolver.Diagnostics, r.Diagnostics) { + t.Errorf("diagnostics after linking %v, want %v", resolver.Diagnostics, r.Diagnostics) + } + } +} + +// A prefix carrying a body resolves its type from the annotated declaration's own +// scope; the linker does the same, so a type visible only there is found. +func TestLinkMetadataBodiesResolvesAPrefixBodyFromTheAnnotatedScope(t *testing.T) { + const name = "prefixed.sysml" + const src = `package P { + metadata def M { attribute n; } + part def C { alias Local for M; } + @M { n = 1; } +}` + build := func() (*symbols.Index, *ast.RootNamespace, *resolve.Resolver) { + p := parser.New(source.New(name, []byte(src))) + root := p.ParseFile() + pkg := root.Members[0].(*ast.Membership).Member.(*ast.Package) + c := pkg.Members[1].(*ast.Membership).Member.(*ast.Definition) + usage := pkg.Members[2].(*ast.Membership).Member.(*ast.PrefixMetadata) + prefix := &ast.PrefixMetadata{Body: usage.Body, HasBody: true, + Type: &ast.QualifiedName{Parts: []ast.NameSegment{{Text: "Local"}}}} + c.Prefixes = []*ast.PrefixMetadata{prefix} + pkg.Members = pkg.Members[:2] + idx := symbols.NewIndexFromDoc(name, root) + r := resolve.New(idx) + r.SetModel(semantics.NewModel(r)) + return idx, root, r + } + idx, root, r := build() + r.ResolveDocument(name, root) + want := ownersOf(idx, idx.DocumentRoot(name)) + pkgScope := idx.DocumentRoot(name).Children()[0] + c := pkgScope.Node().(*ast.Package).Members[1].(*ast.Membership).Member.(*ast.Definition) + body := pkgScope.ChildFor(c.Prefixes[0]) + if body == nil || body.Owner() == nil || idx.GetFQN(body.Owner()) != "P::M" { + t.Fatalf("resolving does not own the prefix body by P::M: %v", body) + } + linkedIdx, _, linker := build() + linker.LinkMetadataBodies(name) + if got := ownersOf(linkedIdx, linkedIdx.DocumentRoot(name)); !reflect.DeepEqual(got, want) { + t.Errorf("linked owners %q, want %q", got, want) + } +} + +// Linking reaches every annotation body of the document: only the one typed by a +// name that does not resolve is left unowned. +func TestLinkMetadataBodiesReachesEveryBody(t *testing.T) { + const name = "linked.sysml" + idx, _, linker := indexedDoc(t, name, linkedMetadataBodies) + linker.LinkMetadataBodies(name) + var bodies, owned int + var visit func(*symbols.Scope) + visit = func(s *symbols.Scope) { + if _, ok := s.Node().(*ast.PrefixMetadata); ok { + bodies++ + if s.Owner() != nil { + owned++ + if fqn := idx.GetFQN(s.Owner()); fqn != "P::M" && fqn != "P::M::a" { + t.Errorf("body owned by %s, want P::M or P::M::a", fqn) + } + } + } + for _, c := range s.Children() { + visit(c) + } + } + visit(idx.DocumentRoot(name)) + if bodies != 5 || owned != 4 { + t.Errorf("%d bodies, %d owned; want 5 and 4", bodies, owned) + } +} + +// parsedRoot parses src as name, failing the test on a parse diagnostic. +func parsedRoot(t *testing.T, name, src string) *ast.RootNamespace { + t.Helper() + p := parser.New(source.New(name, []byte(src))) + root := p.ParseFile() + if len(p.Diagnostics) != 0 { + t.Fatalf("%s: parse diagnostics: %v", name, p.Diagnostics) + } + return root +} + +// An annotation body owned by a definition of another document follows that +// document: once it is replaced, resolving or linking the unchanged annotating +// document owns the body by the definition now indexed, and once the definition +// is gone the body is owned by nothing. +func TestMetadataBodyOwnerFollowsTheDefinitionsDocument(t *testing.T) { + const meta, use = "meta.sysml", "use.sysml" + const useSrc = "package Use { private import Meta::*; part def C { @M { a = b; } } }" + bodyOwner := func(idx *symbols.Index) *symbols.Symbol { + body := idx.DocumentRoot(use).Children()[0].Children()[0].Children()[0] + if _, ok := body.Node().(*ast.PrefixMetadata); !ok { + t.Fatalf("C's first scope is a %T, want the annotation body", body.Node()) + } + return body.Owner() + } + for _, relink := range []struct { + name string + do func(r *resolve.Resolver, root *ast.RootNamespace) + }{ + {"resolving", func(r *resolve.Resolver, root *ast.RootNamespace) { r.ResolveDocument(use, root) }}, + {"linking", func(r *resolve.Resolver, _ *ast.RootNamespace) { r.LinkMetadataBodies(use) }}, + } { + idx := symbols.NewIndex() + idx.AddDocument(meta, parsedRoot(t, meta, "package Meta { metadata def M { attribute a; attribute b; } }")) + useRoot := parsedRoot(t, use, useSrc) + idx.AddDocument(use, useRoot) + idx.ExpandWildcardImports() + r := resolve.New(idx) + r.SetModel(semantics.NewModel(r)) + r.ResolveDocument(use, useRoot) + first := bodyOwner(idx) + if first == nil || idx.GetFQN(first) != "Meta::M" { + t.Fatalf("%s: body owned by %v before the reload, want Meta::M", relink.name, first) + } + + idx.AddDocument(meta, parsedRoot(t, meta, "package Meta { metadata def M { attribute a; attribute c; } }")) + idx.ExpandWildcardImports() + r = resolve.New(idx) + r.SetModel(semantics.NewModel(r)) + relink.do(r, useRoot) + if got := bodyOwner(idx); got == nil || idx.GetFQN(got) != "Meta::M" || got == first { + t.Errorf("%s after the definition's document was replaced: body owned by %s, want the Meta::M now indexed", relink.name, fqnOrNone(idx, got)) + } + if syms := idx.LookupQualified("Meta::M::b"); len(syms) != 0 { + t.Fatalf("Meta::M::b should be gone from the index, found %d", len(syms)) + } + + idx.RemoveDocument(meta) + r = resolve.New(idx) + r.SetModel(semantics.NewModel(r)) + relink.do(r, useRoot) + if got := bodyOwner(idx); got != nil { + t.Errorf("%s after the definition's document was removed: body owned by %s, want nothing", relink.name, fqnOrNone(idx, got)) + } + } +} + +func fqnOrNone(idx *symbols.Index, sym *symbols.Symbol) string { + if sym == nil { + return "nothing" + } + if fqn := idx.GetFQN(sym); fqn != "" { + return fqn + } + return sym.Name + " of an earlier " + sym.DocName +} diff --git a/tests/stressmodel/bench_test.go b/tests/stressmodel/bench_test.go index eb0a37a7fe..1070f84501 100644 --- a/tests/stressmodel/bench_test.go +++ b/tests/stressmodel/bench_test.go @@ -6,7 +6,12 @@ import ( "strings" "testing" + "github.com/Open-MBEE/OpenSysML/internal/check/passes" + "github.com/Open-MBEE/OpenSysML/internal/check/passes/identity" "github.com/Open-MBEE/OpenSysML/internal/frontend/repl" + "github.com/Open-MBEE/OpenSysML/internal/syntax/ast" + "github.com/Open-MBEE/OpenSysML/internal/syntax/parser" + "github.com/Open-MBEE/OpenSysML/internal/syntax/source" "github.com/Open-MBEE/OpenSysML/internal/workspace/model" ) @@ -115,6 +120,90 @@ func BenchmarkEditBeside(b *testing.B) { } } +// BenchmarkValidateSplit loads a network split one file per plane as the command +// line does, on one worker and on one per CPU. +func BenchmarkValidateSplit(b *testing.B) { + for _, n := range networkSizes { + files, stats := splitFiles(network(n)) + for _, jobs := range []int{1, runtime.GOMAXPROCS(0)} { + b.Run(fmt.Sprintf("satellites=%d/files=%d/jobs=%d", stats.Satellites, len(files), jobs), func(b *testing.B) { + load := func() { + sess := repl.NewSession() + if err := sess.SetJobs(jobs); err != nil { + b.Fatal(err) + } + sess.SubmitFiles(files) + if sess.HasErrors() { + b.Fatalf("the split network did not analyse cleanly:\n%s", strings.Join(sess.DiagnosticLines(), "\n")) + } + } + load() + b.ReportAllocs() + b.ResetTimer() + for i := 0; i < b.N; i++ { + load() + } + }) + } + } +} + +// perDocumentRegistry is the default registry without the workspace-wide audits. +func perDocumentRegistry() *passes.Registry { + reg := passes.NewRegistry() + for _, p := range passes.DefaultRegistry().Passes() { + switch p.(type) { + case passes.OOSEMMethodPass, identity.MetadataPass, passes.MOSAPass: + continue + } + reg.Register(p) + } + return reg +} + +// BenchmarkAnalyzeSplitPerDocument analyzes the split network's files over one +// index without the workspace-wide audits: the pool's own speedup. +func BenchmarkAnalyzeSplitPerDocument(b *testing.B) { + for _, n := range networkSizes { + files, stats := splitFiles(network(n)) + idx, _ := model.NewIndexWithStdlib() + roots := make([]*ast.RootNamespace, len(files)) + names := make([]string, len(files)) + for i, f := range files { + p := parser.New(source.New(f.Name, []byte(f.Text))) + roots[i] = p.ParseFile() + if len(p.Diagnostics) > 0 { + b.Fatalf("%s: %s", f.Name, p.Diagnostics[0].Message) + } + names[i] = f.Name + idx.AddDocument(f.Name, roots[i]) + } + idx.ExpandWildcardImports() + batch := &passes.Batch{Documents: names} + passes.PrepareBatch(idx, batch) + reg := perDocumentRegistry() + for _, jobs := range []int{1, runtime.GOMAXPROCS(0)} { + b.Run(fmt.Sprintf("satellites=%d/files=%d/jobs=%d", stats.Satellites, len(files), jobs), func(b *testing.B) { + analyze := func() { + model.ParallelFor(jobs, len(files), func(i int) { + ctx := passes.NewContext(names[i], idx, nil) + ctx.Batch = batch + if diags := reg.Run(ctx, names[i], roots[i]); len(diags) > 0 { + b.Errorf("%s: %s", names[i], diags[0].Message) + } + }) + } + analyze() + b.ReportAllocs() + b.ResetTimer() + for i := 0; i < b.N; i++ { + analyze() + } + }) + } + } +} + // openFiles opens every file of a network in a workspace and analyzes each, // failing on any diagnostic. func openFiles(tb testing.TB, ws *model.Workspace, files []File) { diff --git a/tests/stressmodel/satnet_test.go b/tests/stressmodel/satnet_test.go index 32c786dde7..aa4326c99e 100644 --- a/tests/stressmodel/satnet_test.go +++ b/tests/stressmodel/satnet_test.go @@ -61,6 +61,53 @@ func TestSatelliteNetworkScales(t *testing.T) { } } +// splitFiles is a network split by plane as the CLI loads it, one source per file. +func splitFiles(n SatelliteNetwork) ([]repl.SourceFile, Stats) { + files, stats := n.Split() + srcs := make([]repl.SourceFile, len(files)) + for i, f := range files { + srcs[i] = repl.SourceFile{Name: f.Name, Text: f.Source} + } + return srcs, stats +} + +// TestSatelliteNetworkSplitValidates: the split declares the single file's +// network, loads clean at one worker and at several, and satisfies across files. +func TestSatelliteNetworkSplitValidates(t *testing.T) { + n := SatelliteNetwork{Planes: 2, Satellites: 2, GroundStations: 1} + _, whole := n.Source() + files, stats := splitFiles(n) + if len(files) != n.Planes+2 { + t.Fatalf("got %d files, want one per plane beside the library and the constellation", len(files)) + } + // The split declares one package per plane over the single file's elements. + whole.Bytes, stats.Bytes = 0, 0 + whole.Elements += n.Planes + if stats != whole { + t.Errorf("split declares %+v, the single file %+v", stats, whole) + } + + for _, jobs := range []int{1, 4} { + s := repl.NewSession() + s.SetConformanceMode(diag.ConformanceModeOf(true)) + if err := s.SetJobs(jobs); err != nil { + t.Fatal(err) + } + for _, d := range s.SubmitFiles(files).Diagnostics { + t.Errorf("jobs=%d: diagnostic: %s", jobs, d.Message) + } + verdicts := s.CheckSatisfy("") + if len(verdicts) != stats.Requirements { + t.Fatalf("jobs=%d: got %d satisfy verdicts, want %d", jobs, len(verdicts), stats.Requirements) + } + for _, v := range verdicts { + if !v.Holds() { + t.Errorf("jobs=%d: %s: %v", jobs, v.Subject, v.Lines) + } + } + } +} + // TestSatelliteNetworkFilesValidate keeps the multi-file form in step with the // single one: the same constellation split by plane loads clean under strict // conformance, one document per file, and every satisfy assertion holds. diff --git a/tools/cmd/stress-model/main.go b/tools/cmd/stress-model/main.go index 577e703f2f..b89aee40ac 100644 --- a/tools/cmd/stress-model/main.go +++ b/tools/cmd/stress-model/main.go @@ -1,13 +1,17 @@ -// Command stress-model writes a large generated SysML v2 model of a stated shape -// and size to stdout, for measuring how the toolchain scales. The one shape today -// is a satellite network — a constellation of fully modeled spacecraft and the -// ground stations they downlink to. See docs/project/satellite-network-stress-test.md. +// Command stress-model writes a large generated satellite-network model to stdout, +// or one file per orbital plane; see docs/project/satellite-network-stress-test.md. package main import ( + "crypto/sha256" + "encoding/hex" + "errors" "flag" "fmt" "os" + "path/filepath" + "slices" + "strings" "github.com/Open-MBEE/OpenSysML/tests/stressmodel" ) @@ -17,6 +21,7 @@ func main() { perPlane := flag.Int("satellites", 8, "satellites in each orbital plane") stations := flag.Int("ground-stations", 3, "ground stations the constellation downlinks to") stats := flag.Bool("stats", false, "report on stderr what the model declares") + split := flag.String("split-planes", "", "write one .sysml per orbital plane, beside the library and the constellation, into this directory instead of stdout") flag.Parse() if flag.NArg() != 0 { fmt.Fprintf(os.Stderr, "stress-model: unexpected argument %q\n", flag.Arg(0)) @@ -27,7 +32,15 @@ func main() { os.Exit(2) } n := stressmodel.SatelliteNetwork{Planes: *planes, Satellites: *perPlane, GroundStations: *stations} - s, err := n.Generate(os.Stdout) + var ( + s stressmodel.Stats + err error + ) + if *split != "" { + s, err = writeSplit(n, *split) + } else { + s, err = n.Generate(os.Stdout) + } if err != nil { fmt.Fprintf(os.Stderr, "stress-model: %v\n", err) os.Exit(1) @@ -37,3 +50,193 @@ func main() { s.Satellites, s.GroundStations, s.Components, s.Connections, s.Requirements, s.Elements, s.Bytes) } } + +// manifestName is the file in a -split-planes directory listing what the last +// generation wrote there, each file with the digest of its content, so the +// next one removes only its own files and only while they read as written. +const manifestName = ".stress-model-files" + +// digest is the manifest's identity of a file's content. +func digest(content []byte) string { + sum := sha256.Sum256(content) + return hex.EncodeToString(sum[:]) +} + +// writeSplit writes the network one file per plane into dir, creating it, and +// removes what an earlier generation wrote there that this one did not. +func writeSplit(n stressmodel.SatelliteNetwork, dir string) (stressmodel.Stats, error) { + files, stats := n.Split() + if err := os.MkdirAll(dir, 0o750); err != nil { + return stats, err + } + previous, err := readManifest(dir) + if err != nil { + return stats, err + } + if err := replaceable(dir, files, previous); err != nil { + return stats, err + } + staging, err := os.MkdirTemp(dir, ".stress-model-*") + if err != nil { + return stats, err + } + defer os.RemoveAll(staging) + for _, f := range files { + if err := os.WriteFile(filepath.Join(staging, f.Name), []byte(f.Source), 0o600); err != nil { + return stats, err + } + } + // Recorded before any move, so a failed generation leaves nothing unrecorded. + if err := replaceManifest(dir, staging, intended(previous, files)); err != nil { + return stats, err + } + written := make(map[string]bool, len(files)) + for _, f := range files { + if err := os.Rename(filepath.Join(staging, f.Name), filepath.Join(dir, f.Name)); err != nil { + return stats, err + } + written[f.Name] = true + } + for _, rec := range previous { + if written[rec.Name] { + continue + } + if err := removeIfRecorded(filepath.Join(dir, rec.Name), rec.Digest); err != nil { + return stats, err + } + } + return stats, replaceManifest(dir, staging, manifestOf(files)) +} + +// intended is the manifest that stands while a generation moves its files in: +// the last generation's records, then a record of each file not already among them. +func intended(previous []record, files []stressmodel.File) string { + var b strings.Builder + standing := make(map[record]bool, len(previous)) + for _, rec := range previous { + standing[rec] = true + b.WriteString(rec.Digest + " " + rec.Name + "\n") + } + for _, f := range files { + if rec := (record{Name: f.Name, Digest: digest([]byte(f.Source))}); !standing[rec] { + b.WriteString(rec.Digest + " " + rec.Name + "\n") + } + } + return b.String() +} + +// manifestOf records each of files with the digest of its content. +func manifestOf(files []stressmodel.File) string { + var b strings.Builder + for _, f := range files { + b.WriteString(digest([]byte(f.Source)) + " " + f.Name + "\n") + } + return b.String() +} + +// replaceManifest puts content in place as the manifest in one rename, so a +// manifest that stands is never cut short or half written. +func replaceManifest(dir, staging, content string) error { + next := filepath.Join(staging, manifestName) + if err := os.WriteFile(next, []byte(content), 0o600); err != nil { + return err + } + return os.Rename(next, filepath.Join(dir, manifestName)) +} + +// replaceable reports an error naming every file the generation would write +// over that the last generation did not write, or that has changed since: the +// generator replaces only its own unedited output, and writes nothing otherwise. +func replaceable(dir string, files []stressmodel.File, previous []record) error { + recorded := make(map[string][]string, len(previous)) + for _, rec := range previous { + recorded[rec.Name] = append(recorded[rec.Name], rec.Digest) + } + var errs []error + for _, f := range files { + path := filepath.Join(dir, f.Name) + info, err := os.Lstat(path) + if errors.Is(err, os.ErrNotExist) { + continue + } + if err != nil { + return err + } + want, ok := recorded[f.Name] + switch { + case !info.Mode().IsRegular(): + errs = append(errs, fmt.Errorf("%s: not a regular file", path)) + case !ok: + errs = append(errs, fmt.Errorf("%s: not written by the last generation into %s", path, dir)) + default: + content, err := os.ReadFile(path) // #nosec G304 -- path is a generated name under the output directory. + if err != nil { + return err + } + if !slices.Contains(want, digest(content)) { + errs = append(errs, fmt.Errorf("%s: changed since the last generation wrote it", path)) + } + } + } + if len(errs) == 0 { + return nil + } + return fmt.Errorf("nothing written: %w", errors.Join(errs...)) +} + +// removeIfRecorded removes the regular file at path when its content still +// has the recorded digest; anything else standing there is not the generator's. +func removeIfRecorded(path, recorded string) error { + info, err := os.Lstat(path) + if errors.Is(err, os.ErrNotExist) { + return nil + } + if err != nil { + return err + } + if !info.Mode().IsRegular() { + return nil + } + content, err := os.ReadFile(path) // #nosec G304 -- path is a recorded name under the output directory. + if errors.Is(err, os.ErrNotExist) { + return nil + } + if err != nil { + return err + } + if digest(content) != recorded { + return nil + } + if err := os.Remove(path); err != nil && !errors.Is(err, os.ErrNotExist) { + return err + } + return nil +} + +// record is one manifest line: a file the last generation wrote, and the +// digest of what it wrote there. +type record struct { + Name string + Digest string +} + +// readManifest returns what the last generation into dir recorded; nothing +// when there was no generation. Only plain names in dir with a digest are honored. +func readManifest(dir string) ([]record, error) { + data, err := os.ReadFile(filepath.Join(dir, manifestName)) // #nosec G304 -- the output directory is named on the command line. + if errors.Is(err, os.ErrNotExist) { + return nil, nil + } + if err != nil { + return nil, err + } + var records []record + for _, line := range strings.Split(string(data), "\n") { + sum, name, ok := strings.Cut(line, " ") + if !ok || sum == "" || name == "" || name == manifestName || filepath.Base(name) != name { + continue + } + records = append(records, record{Name: name, Digest: sum}) + } + return records, nil +} diff --git a/tools/cmd/stress-model/main_test.go b/tools/cmd/stress-model/main_test.go new file mode 100644 index 0000000000..d9e7ce604c --- /dev/null +++ b/tools/cmd/stress-model/main_test.go @@ -0,0 +1,318 @@ +package main + +import ( + "io" + "os" + "path/filepath" + "runtime" + "slices" + "strings" + "testing" + + "github.com/Open-MBEE/OpenSysML/tests/stressmodel" +) + +func TestWriteSplitDropsOnlyWhatALargerGenerationWrote(t *testing.T) { + dir := t.TempDir() + // A model of the user's own, one of them under a name a plane could take. + for _, name := range []string{"notes.sysml", "plane009.sysml"} { + if err := os.WriteFile(filepath.Join(dir, name), []byte("package Notes;\n"), 0o600); err != nil { + t.Fatal(err) + } + } + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 8, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + // A directory now standing where the first generation put a plane. + if err := os.Remove(filepath.Join(dir, "plane007.sysml")); err != nil { + t.Fatal(err) + } + if err := os.Mkdir(filepath.Join(dir, "plane007.sysml"), 0o750); err != nil { + t.Fatal(err) + } + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 4, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + want := []string{ + manifestName, "constellation.sysml", "library.sysml", "notes.sysml", + "plane000.sysml", "plane001.sysml", "plane002.sysml", "plane003.sysml", "plane007.sysml", "plane009.sysml", + } + if got := listing(t, dir); !slices.Equal(got, want) { + t.Errorf("after regenerating with four planes the directory holds %v, want %v", got, want) + } + if got, err := os.ReadFile(filepath.Join(dir, "plane009.sysml")); err != nil || string(got) != "package Notes;\n" { + t.Errorf("the user's plane009.sysml reads %q, %v; want it untouched", got, err) + } +} + +// A first generation into a directory holding files at generated names writes +// nothing: it names every one of them, leaves each as it was and makes no manifest. +func TestWriteSplitWritesNothingOverFilesItDidNotWrite(t *testing.T) { + dir := t.TempDir() + for _, name := range []string{"plane000.sysml", "plane005.sysml", "library.sysml"} { + if err := os.WriteFile(filepath.Join(dir, name), []byte("package Mine;\n"), 0o600); err != nil { + t.Fatal(err) + } + } + _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 2, Satellites: 1}, dir) + if err == nil { + t.Fatal("generating over the user's library.sysml and plane000.sysml should fail") + } + for _, name := range []string{"library.sysml", "plane000.sysml"} { + if !strings.Contains(err.Error(), name) { + t.Errorf("the error %q does not name %s", err, name) + } + } + if strings.Contains(err.Error(), "plane005") { + t.Errorf("the error %q names plane005.sysml, which no generation of two planes writes", err) + } + if got := listing(t, dir); !slices.Equal(got, []string{"library.sysml", "plane000.sysml", "plane005.sysml"}) { + t.Errorf("the refused generation left %v; want the user's three files alone", got) + } + for _, name := range []string{"plane000.sysml", "plane005.sysml", "library.sysml"} { + if got, err := os.ReadFile(filepath.Join(dir, name)); err != nil || string(got) != "package Mine;\n" { + t.Errorf("%s reads %q, %v; want it untouched", name, got, err) + } + } +} + +// A later generation replaces only what the last one wrote and nothing has +// changed since: a plane the user edited, a directory at a plane's name and a +// file of the user's at one all stop it before it writes a byte, and are named +// together; with them out of the way the same generation goes through. +func TestWriteSplitReplacesOnlyItsOwnUnchangedOutput(t *testing.T) { + dir := t.TempDir() + if err := os.WriteFile(filepath.Join(dir, "notes.sysml"), []byte("package Notes;\n"), 0o600); err != nil { + t.Fatal(err) + } + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 2, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + firstListing, firstManifest := listing(t, dir), recordedNames(t, dir) + if err := os.WriteFile(filepath.Join(dir, "plane001.sysml"), []byte("package Edited;\n"), 0o600); err != nil { + t.Fatal(err) + } + if err := os.Mkdir(filepath.Join(dir, "plane002.sysml"), 0o750); err != nil { + t.Fatal(err) + } + if err := os.WriteFile(filepath.Join(dir, "plane003.sysml"), []byte("package Mine;\n"), 0o600); err != nil { + t.Fatal(err) + } + before := append(slices.Clone(firstListing), "plane002.sysml", "plane003.sysml") + slices.Sort(before) + _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 4, Satellites: 1}, dir) + if err == nil { + t.Fatal("generating over an edited plane, a directory and the user's file should fail") + } + for _, name := range []string{"plane001.sysml", "plane002.sysml", "plane003.sysml"} { + if !strings.Contains(err.Error(), name) { + t.Errorf("the error %q does not name %s", err, name) + } + } + if got := listing(t, dir); !slices.Equal(got, before) { + t.Errorf("the refused generation left %v, want %v: nothing written, nothing staged", got, before) + } + if got, err := os.ReadFile(filepath.Join(dir, "plane001.sysml")); err != nil || string(got) != "package Edited;\n" { + t.Errorf("the edited plane001.sysml reads %q, %v; want the edit kept", got, err) + } + if got := recordedNames(t, dir); !slices.Equal(got, firstManifest) { + t.Errorf("the manifest reads %v after the refused generation, want %v as before", got, firstManifest) + } + + for _, name := range []string{"plane001.sysml", "plane002.sysml", "plane003.sysml"} { + if err := os.RemoveAll(filepath.Join(dir, name)); err != nil { + t.Fatal(err) + } + } + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 4, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + want := []string{manifestName, "constellation.sysml", "library.sysml", "notes.sysml", "plane000.sysml", "plane001.sysml", "plane002.sysml", "plane003.sysml"} + if got := listing(t, dir); !slices.Equal(got, want) { + t.Errorf("regenerating with four planes left %v, want %v", got, want) + } +} + +// A generation interrupted between recording a file and moving it in leaves +// the manifest naming whatever stood at that name; a later generation must +// not take that for its own. Neither may it take a recorded file the user +// has since edited. +func TestWriteSplitRemovesOnlyFilesThatReadAsRecorded(t *testing.T) { + dir := t.TempDir() + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 3, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + // The user's plane007.sysml, recorded by an interrupted larger generation + // that never moved its own plane007 over it. + if err := os.WriteFile(filepath.Join(dir, "plane007.sysml"), []byte("package Mine;\n"), 0o600); err != nil { + t.Fatal(err) + } + interrupted := digest([]byte("package Plane7;\n")) + " plane007.sysml\n" + manifest, err := os.OpenFile(filepath.Join(dir, manifestName), os.O_WRONLY|os.O_APPEND, 0o600) + if err != nil { + t.Fatal(err) + } + if _, err := manifest.WriteString(interrupted); err != nil { + t.Fatal(err) + } + if err := manifest.Close(); err != nil { + t.Fatal(err) + } + // The user's edit of a plane the generator did write. + if err := os.WriteFile(filepath.Join(dir, "plane002.sysml"), []byte("package Edited;\n"), 0o600); err != nil { + t.Fatal(err) + } + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 1, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + want := []string{manifestName, "constellation.sysml", "library.sysml", "plane000.sysml", "plane002.sysml", "plane007.sysml"} + if got := listing(t, dir); !slices.Equal(got, want) { + t.Errorf("regenerating with one plane left %v, want %v: plane001 removed, the user's files kept", got, want) + } + for name, content := range map[string]string{"plane002.sysml": "package Edited;\n", "plane007.sysml": "package Mine;\n"} { + if got, err := os.ReadFile(filepath.Join(dir, name)); err != nil || string(got) != content { + t.Errorf("%s reads %q, %v; want it untouched", name, got, err) + } + } + if recorded := recordedNames(t, dir); !slices.Equal(recorded, []string{"library.sysml", "plane000.sysml", "constellation.sysml"}) { + t.Errorf("the manifest reads %v; want only the last generation's files", recorded) + } +} + +// After a generation interrupted between recording and moving its files in, +// names hold the new content or the old; the next run replaces both. +func TestWriteSplitRetriesAnInterruptedGeneration(t *testing.T) { + dir := t.TempDir() + first := stressmodel.SatelliteNetwork{Planes: 2, Satellites: 1} + second := stressmodel.SatelliteNetwork{Planes: 2, Satellites: 2} + if _, err := writeSplit(first, dir); err != nil { + t.Fatal(err) + } + files, _ := second.Split() + if files[1].Name != "plane000.sysml" || files[2].Name != "plane001.sysml" { + t.Fatalf("the split writes %s, %s second and third; the test wants the two planes", files[1].Name, files[2].Name) + } + // The second generation recorded both planes and moved only the first in. + if err := os.WriteFile(filepath.Join(dir, files[1].Name), []byte(files[1].Source), 0o600); err != nil { + t.Fatal(err) + } + if err := interruptManifest(dir, files[1:3]); err != nil { + t.Fatal(err) + } + if _, err := writeSplit(second, dir); err != nil { + t.Fatalf("the interrupted generation cannot be retried: %v", err) + } + for _, f := range files { + if got, err := os.ReadFile(filepath.Join(dir, f.Name)); err != nil || string(got) != f.Source { + t.Errorf("%s does not read as the retried generation writes it (%v)", f.Name, err) + } + } + if recorded := recordedNames(t, dir); !slices.Equal(recorded, []string{"library.sysml", "plane000.sysml", "plane001.sysml", "constellation.sysml"}) { + t.Errorf("the manifest reads %v; want each of the retried generation's files once", recorded) + } +} + +// A file an interrupted generation recorded and moved in is still the +// generator's: a smaller generation afterwards removes it like any other. +func TestWriteSplitRemovesWhatAnInterruptedGenerationMovedIn(t *testing.T) { + dir := t.TempDir() + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 2, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + files, _ := stressmodel.SatelliteNetwork{Planes: 2, Satellites: 2}.Split() + if err := os.WriteFile(filepath.Join(dir, files[2].Name), []byte(files[2].Source), 0o600); err != nil { + t.Fatal(err) + } + if err := interruptManifest(dir, files[1:3]); err != nil { + t.Fatal(err) + } + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 1, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + want := []string{manifestName, "constellation.sysml", "library.sysml", "plane000.sysml"} + if got := listing(t, dir); !slices.Equal(got, want) { + t.Errorf("regenerating with one plane left %v, want %v", got, want) + } +} + +// interruptManifest records files in dir's manifest the way a generation does +// before moving them in, and stops there. +func interruptManifest(dir string, files []stressmodel.File) error { + previous, err := readManifest(dir) + if err != nil { + return err + } + staging, err := os.MkdirTemp(dir, ".stress-model-*") + if err != nil { + return err + } + defer os.RemoveAll(staging) + return replaceManifest(dir, staging, intended(previous, files)) +} + +// The manifest that stood while a generation ran is replaced whole, never +// cut short: a reader holding it sees it unchanged. +func TestWriteSplitReplacesTheManifestWholeInsteadOfTruncatingIt(t *testing.T) { + if runtime.GOOS == "windows" { + t.Skip("a renamed-over file cannot be held open on Windows") + } + dir := t.TempDir() + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 2, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + first, err := os.ReadFile(filepath.Join(dir, manifestName)) + if err != nil { + t.Fatal(err) + } + old, err := os.Open(filepath.Join(dir, manifestName)) + if err != nil { + t.Fatal(err) + } + defer old.Close() + if _, err := writeSplit(stressmodel.SatelliteNetwork{Planes: 1, Satellites: 1}, dir); err != nil { + t.Fatal(err) + } + second, err := os.ReadFile(filepath.Join(dir, manifestName)) + if err != nil { + t.Fatal(err) + } + held, err := io.ReadAll(old) + if err != nil { + t.Fatal(err) + } + if string(held) != string(first) { + t.Errorf("the manifest held open through the second generation reads:\n%s\nwant the first manifest unchanged:\n%s", held, first) + } + if string(second) == string(first) { + t.Error("the second generation did not replace the manifest") + } + if recorded := recordedNames(t, dir); !slices.Equal(recorded, []string{"library.sysml", "plane000.sysml", "constellation.sysml"}) { + t.Errorf("the manifest reads %v; want only the second generation's files", recorded) + } +} + +func recordedNames(t *testing.T, dir string) []string { + t.Helper() + records, err := readManifest(dir) + if err != nil { + t.Fatal(err) + } + var names []string + for _, r := range records { + names = append(names, r.Name) + } + return names +} + +func listing(t *testing.T, dir string) []string { + t.Helper() + entries, err := os.ReadDir(dir) + if err != nil { + t.Fatal(err) + } + var names []string + for _, e := range entries { + names = append(names, e.Name()) + } + return names +}