Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ bin/package-extension.sh # Package Chrome extension into build
Tests (run in CI on every push and PR — see `.github/workflows/test.yml`):

```bash
./test/smoke-test.sh # 401 assertions: install, SKILL.md freshness, frontmatter, version agreement,
./test/smoke-test.sh # 405 assertions: install, SKILL.md freshness, frontmatter, version agreement,
# canonical section names, /idstack: namespacing, resolve-snippet lockstep, bash -n,
# Claude-Code-only invariant (no dist/, no AGENTS.md, no retired-CLI references)
./test/integration-test.sh # 51 behavioral tests across the bin/ scripts; also proves the suite
Expand All @@ -48,6 +48,7 @@ Tests (run in CI on every push and PR — see `.github/workflows/test.yml`):
python3 test/test-ste-check.py # bin/idstack-ste-check unit tests (smoke-test also runs them)
python3 test/check-evidence-cards.py . # Verifies landing page evidence cards match evidence/references.md
python3 test/check-citation-tiers.py . # Verifies each [Code-N] [Tn] in skills, templates, README and landing page has its references.md tier
python3 test/test-citation-tiers.py # check-citation-tiers.py unit tests: comma lists, line-split citations (smoke-test also runs them)
python3 test/check-doc-accuracy.py . # Verifies documentation accuracy across version strings, binaries, flags, and links
./test/mutation-test.sh # Reintroduces each fixed defect and asserts its guarding test fails
```
Expand Down
1 change: 1 addition & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,6 +147,7 @@ Every suite below runs in CI (`.github/workflows/test.yml`) on every push and pu
| `test/test-responsive-landing.js` | Responsive and mobile-ergonomics invariants for `docs/index.html` — fluid tokens, notch-safe gutters, breakpoint-scoped rules, touch targets. Runs on node, via `smoke-test.sh` |
| `python3 test/check-evidence-cards.py .` | Verifies landing page evidence card study counts and tier ranges against `evidence/references.md` |
| `python3 test/check-citation-tiers.py .` | Verifies that each `[Code-N] [Tn]` citation in the skill templates, `templates/`, README and landing page states the tier that `evidence/references.md` gives that code. `test/smoke-test.sh` runs it |
| `python3 test/test-citation-tiers.py` | `test/check-citation-tiers.py` on small test repos: codes with commas between them, a citation that a line break splits, and a blank line that ends a citation. `test/smoke-test.sh` runs it |
| `python3 test/test-ste-check.py` | `bin/idstack-ste-check`: each rule that it examines, the text that it does not examine, and the clean fixtures in `test/fixtures/ste/`. `test/smoke-test.sh` runs it |
| `python3 test/check-doc-accuracy.py .` | Validates version agreement, manifest schema version, binary/flag references, link targets, and surface accuracy across docs |
| `test/mutation-test.sh` | Reintroduces each known defect into a throwaway copy and asserts the guarding test fails. Add a mutation here whenever you fix a bug — it is what proves your new test would have caught it |
Expand Down
2 changes: 1 addition & 1 deletion skills/course-import/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -1153,7 +1153,7 @@ IFS= read -r _PROJECT_NAME <<'IDSTACK_PROJECT_NAME'
<project name>
IDSTACK_PROJECT_NAME
_SLUG=$("$_IDSTACK/bin/idstack-slugify" "$_PROJECT_NAME" 2>/dev/null || echo "untitled-course")
[ "$_SLUG" = "project-name" ] && echo "PROJECT_NAME_NOT_SET: replace <project name> with the course title and run this command again."
[ "$_SLUG" = "project-name" ] && { echo "PROJECT_NAME_NOT_SET: replace <project name> with the course title and run this command again."; exit 1; }
_EXPORT_DIR=".idstack/exports/$_SLUG"
_REPORT_PATH="$_EXPORT_DIR/course-import.html"
mkdir -p "$_EXPORT_DIR/assets"
Expand Down
2 changes: 1 addition & 1 deletion skills/course-import/SKILL.md.tmpl
Original file line number Diff line number Diff line change
Expand Up @@ -721,7 +721,7 @@ IFS= read -r _PROJECT_NAME <<'IDSTACK_PROJECT_NAME'
<project name>
IDSTACK_PROJECT_NAME
_SLUG=$("$_IDSTACK/bin/idstack-slugify" "$_PROJECT_NAME" 2>/dev/null || echo "untitled-course")
[ "$_SLUG" = "project-name" ] && echo "PROJECT_NAME_NOT_SET: replace <project name> with the course title and run this command again."
[ "$_SLUG" = "project-name" ] && { echo "PROJECT_NAME_NOT_SET: replace <project name> with the course title and run this command again."; exit 1; }
_EXPORT_DIR=".idstack/exports/$_SLUG"
_REPORT_PATH="$_EXPORT_DIR/course-import.html"
mkdir -p "$_EXPORT_DIR/assets"
Expand Down
2 changes: 1 addition & 1 deletion skills/needs-analysis/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -761,7 +761,7 @@ IFS= read -r _PROJECT_NAME <<'IDSTACK_PROJECT_NAME'
<project name>
IDSTACK_PROJECT_NAME
_SLUG=$("$_IDSTACK/bin/idstack-slugify" "$_PROJECT_NAME" 2>/dev/null || echo "untitled-course")
[ "$_SLUG" = "project-name" ] && echo "PROJECT_NAME_NOT_SET: replace <project name> with the course title and run this command again."
[ "$_SLUG" = "project-name" ] && { echo "PROJECT_NAME_NOT_SET: replace <project name> with the course title and run this command again."; exit 1; }
_EXPORT_DIR=".idstack/exports/$_SLUG"
_REPORT_PATH="$_EXPORT_DIR/needs-analysis.html"
mkdir -p "$_EXPORT_DIR/assets"
Expand Down
2 changes: 1 addition & 1 deletion skills/needs-analysis/SKILL.md.tmpl
Original file line number Diff line number Diff line change
Expand Up @@ -329,7 +329,7 @@ IFS= read -r _PROJECT_NAME <<'IDSTACK_PROJECT_NAME'
<project name>
IDSTACK_PROJECT_NAME
_SLUG=$("$_IDSTACK/bin/idstack-slugify" "$_PROJECT_NAME" 2>/dev/null || echo "untitled-course")
[ "$_SLUG" = "project-name" ] && echo "PROJECT_NAME_NOT_SET: replace <project name> with the course title and run this command again."
[ "$_SLUG" = "project-name" ] && { echo "PROJECT_NAME_NOT_SET: replace <project name> with the course title and run this command again."; exit 1; }
_EXPORT_DIR=".idstack/exports/$_SLUG"
_REPORT_PATH="$_EXPORT_DIR/needs-analysis.html"
mkdir -p "$_EXPORT_DIR/assets"
Expand Down
50 changes: 27 additions & 23 deletions test/check-citation-tiers.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,9 @@
[A-1] [T1] the code must be T1 in references.md
[A-1] [B-2] [T1] each code must be T1 in references.md
[A-1] [B-2] [T1] [T3] refused: put the tier after each code instead
Each [Code-N] must also be an entry in references.md.
Codes can have commas between them ("[A-1], [B-2] [T1]"), and one line break
can come between two items of a citation. A blank line ends it. Each [Code-N] must also be an entry
in references.md.

Scope: skills/*/SKILL.md.tmpl, templates/, README.md and docs/index.html. The
generated SKILL.md files are left out (they repeat the templates, and
Expand All @@ -33,8 +35,10 @@
# "- [CogLoad-4] Sweller, J. (1994). ... *Learning and Instruction*. T5"
REF_RE = re.compile(r"^- \[([A-Za-z]+-\d+)\] .* (T[1-5])\s*$")
CODE_RE = re.compile(r"\[([A-Z][A-Za-z]*-\d+)\]")
# Between two items: spaces of any kind, an optional comma, and at most one line break.
SEP = r"[^\S\n]*(?:,[^\S\n]*)?(?:\n[^\S\n]*)?"
# One or more codes, then one or more tiers.
GROUP_RE = re.compile(r"((?:\[[A-Z][A-Za-z]*-\d+\]\s*)+)((?:\[T[1-5]\]\s*)+)")
GROUP_RE = re.compile(r"((?:\[[A-Z][A-Za-z]*-\d+\]" + SEP + r")+)((?:\[T[1-5]\]" + SEP + r")+)")
# The landing page shows a tier as <span class="tier tier-1">T1</span>.
TIER_SPAN_RE = re.compile(r'<span class="tier tier-[1-5]">(T[1-5])</span>')
TAG_RE = re.compile(r"<[^>]+>")
Expand Down Expand Up @@ -79,29 +83,29 @@ def main():
cited = 0
for path in sources(root):
rel = os.path.relpath(path, root)
lines = io.open(path, encoding="utf-8").read().split("\n")
for n, raw in enumerate(lines, 1):
line = plain(raw)
where = "%s:%d" % (rel, n)
lines = [plain(raw) for raw in io.open(path, encoding="utf-8").read().split("\n")]
found = []
for n, line in enumerate(lines, 1):
for code in CODE_RE.findall(line):
if code not in refs:
problems.append("%s: [%s] is not in evidence/references.md" % (where, code))
for m in GROUP_RE.finditer(line):
codes = CODE_RE.findall(m.group(1))
tiers = re.findall(r"T[1-5]", m.group(2))
cited += 1
if len(tiers) > 1:
problems.append(
"%s: %s gives more than one tier. Put the tier after each code."
% (where, m.group(0).strip())
)
continue
for code in codes:
if code in refs and refs[code] != tiers[0]:
problems.append(
"%s: [%s] is %s in evidence/references.md, not %s"
% (where, code, refs[code], tiers[0])
)
found.append((n, "[%s] is not in evidence/references.md" % code))
# Match on the whole file, so that a citation split by a line break is seen.
text = "\n".join(lines)
for m in GROUP_RE.finditer(text):
n = text.count("\n", 0, m.start()) + 1
codes = CODE_RE.findall(m.group(1))
tiers = re.findall(r"T[1-5]", m.group(2))
cited += 1
if len(tiers) > 1:
found.append((n, "%s gives more than one tier. Put the tier after each code."
% " ".join(m.group(0).split())))
continue
for code in codes:
if code in refs and refs[code] != tiers[0]:
found.append((n, "[%s] is %s in evidence/references.md, not %s"
% (code, refs[code], tiers[0])))
for n, message in sorted(found, key=lambda item: item[0]):
problems.append("%s:%d: %s" % (rel, n, message))

# A pattern that stops matching would make every check above vacuous.
if cited == 0:
Expand Down
61 changes: 61 additions & 0 deletions test/mutation-test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1993,6 +1993,67 @@ PY
regen
expect_fail "course-builder gives two codes two tiers in one group" "$WORK/r/test/smoke-test.sh" "$WORK/r"

# 54a-54d. Small defects that the review of the follow-up fixes found.
# 54a. needs-analysis only warns about an unreplaced course title again -> smoke-test
# must fail. A warning alone left an empty exports/project-name/ folder.
fresh
python3 - "$WORK/r/skills/needs-analysis/SKILL.md.tmpl" <<'PY'
import sys
p = sys.argv[1]; s = open(p, encoding='utf-8').read()
old = ('[ "$_SLUG" = "project-name" ] && { echo "PROJECT_NAME_NOT_SET: replace <project name> '
'with the course title and run this command again."; exit 1; }\n')
assert s.count(old) == 1, 'anchor not unique: %d' % s.count(old)
s = s.replace(old, '[ "$_SLUG" = "project-name" ] && echo "PROJECT_NAME_NOT_SET: replace <project name> '
'with the course title and run this command again."\n', 1)
open(p, 'w', encoding='utf-8').write(s)
PY
regen
expect_fail "needs-analysis makes a folder for an unreplaced course title" "$WORK/r/test/smoke-test.sh" "$WORK/r"

# 54b. course-import only warns about an unreplaced course title again -> smoke-test
# must fail. Each skill has its own case, as in 52a-52b.
fresh
python3 - "$WORK/r/skills/course-import/SKILL.md.tmpl" <<'PY'
import sys
p = sys.argv[1]; s = open(p, encoding='utf-8').read()
old = ('[ "$_SLUG" = "project-name" ] && { echo "PROJECT_NAME_NOT_SET: replace <project name> '
'with the course title and run this command again."; exit 1; }\n')
assert s.count(old) == 1, 'anchor not unique: %d' % s.count(old)
s = s.replace(old, '[ "$_SLUG" = "project-name" ] && echo "PROJECT_NAME_NOT_SET: replace <project name> '
'with the course title and run this command again."\n', 1)
open(p, 'w', encoding='utf-8').write(s)
PY
regen
expect_fail "course-import makes a folder for an unreplaced course title" "$WORK/r/test/smoke-test.sh" "$WORK/r"

# 54c. The citation-tier checker matches only codes in one line with spaces between
# them again -> test-citation-tiers must fail. Then a code before a comma, or a
# citation that a line break splits, gets no check.
fresh
python3 - "$WORK/r/test/check-citation-tiers.py" <<'PY'
import sys
p = sys.argv[1]; s = open(p, encoding='utf-8').read()
old = 'SEP = r"[^\\S\\n]*(?:,[^\\S\\n]*)?(?:\\n[^\\S\\n]*)?"\n'
assert s.count(old) == 1, 'anchor not unique: %d' % s.count(old)
s = s.replace(old, 'SEP = r"[^\\S\\n]*"\n', 1)
open(p, 'w', encoding='utf-8').write(s)
PY
expect_fail "the citation-tier checker misses commas and line breaks" python3 "$WORK/r/test/test-citation-tiers.py"

# 54d. The rendered landing test reports a stopped Chrome before stderr closes again ->
# smoke-test must fail. Then output that Chrome's helpers write after the main
# process exits is lost from the message.
fresh
python3 - "$WORK/r/test/test-rendered-landing.js" <<'PY'
import sys
p = sys.argv[1]; s = open(p, encoding='utf-8').read()
old = " if (chrome.stderr.readableEnded) return fail();\n"
assert s.count(old) == 1, 'anchor not unique: %d' % s.count(old)
s = s.replace(old, " return fail();\n", 1)
open(p, 'w', encoding='utf-8').write(s)
PY
expect_fail "the rendered landing test loses output that Chrome writes after it exits" "$WORK/r/test/smoke-test.sh" "$WORK/r"

echo ""
echo "guarded: $pass NOT guarded: $fail skipped: $skip"
[ "$fail" -eq 0 ]
Expand Down
30 changes: 30 additions & 0 deletions test/smoke-test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -142,6 +142,9 @@ if command -v python3 &>/dev/null; then
$TIER_DRIFT"
check "skill and template citations state their evidence/references.md tier" \
"if [ -n \"\$TIER_DRIFT\" ]; then printf '%s\n' \"\$TIER_DRIFT\"; false; fi"
# A clean repo proves nothing about a form the checker cannot see. These tests
# give it codes separated by commas and a citation that a line break splits.
check "citation-tier checker unit tests pass" "python3 '$IDSTACK_DIR/test/test-citation-tiers.py'"
fi

check "doc accuracy check passes" "python3 '$IDSTACK_DIR/test/check-doc-accuracy.py' '$IDSTACK_DIR'"
Expand Down Expand Up @@ -219,6 +222,24 @@ sys.exit(0 if ok else 1)'
for skill in needs-analysis course-import; do
check "$skill takes the report slug from the course title, not the manifest it writes later" "python3 -c '$SLUG_SOURCE_PY' '$IDSTACK_DIR/skills/$skill/SKILL.md.tmpl'"
done
# If the model runs the slug block with <project name> still in it, the block must
# stop before it makes a folder. A warning alone left an empty exports/project-name/.
# Run the generated block in an empty directory, as a skill does.
SLUG_GUARD_PY='import os, re, shutil, subprocess, sys, tempfile
s = open(sys.argv[1], encoding="utf-8").read()
block = [b for b in re.findall(r"```bash\n(.*?)```", s, re.S) if "idstack-slugify" in b][0]
env = dict(os.environ, IDSTACK_HOME=sys.argv[2])
env.pop("CLAUDE_PLUGIN_ROOT", None)
d = tempfile.mkdtemp()
r = subprocess.run(["bash", "-c", block], cwd=d, env=env, stdout=subprocess.PIPE,
stderr=subprocess.STDOUT, universal_newlines=True)
made = os.path.exists(os.path.join(d, ".idstack", "exports", "project-name"))
shutil.rmtree(d)
print(r.stdout)
sys.exit(0 if "PROJECT_NAME_NOT_SET" in r.stdout and r.returncode != 0 and not made else 1)'
for skill in needs-analysis course-import; do
check "$skill stops before it makes a folder for an unreplaced course title" "python3 -c '$SLUG_GUARD_PY' '$IDSTACK_DIR/skills/$skill/SKILL.md' '$IDSTACK_DIR'"
done

# Canonical manifest section names only — these five non-canonical tokens once
# shipped in re-run checks and prose, making re-run detection dead in 5 skills.
Expand Down Expand Up @@ -426,6 +447,14 @@ if command -v node &>/dev/null; then
check "rendered landing test reports why Chrome stopped at startup" \
"CHROME_PATH='$FAKE_CHROME_DIR/chrome' node '$IDSTACK_DIR/test/test-rendered-landing.js'" \
1 "exit code 3.*fake-chrome-startup-failure"
# Chrome's helper processes can write to stderr after the main process exits. The
# message must still give that output, so the test waits for stderr to close. The
# stub's child writes 0.2s after the stub exits, which makes the case certain.
printf '#!/bin/sh\n(sleep 0.2; echo "fake-chrome-late-output" >&2) &\nexit 4\n' > "$FAKE_CHROME_DIR/late"
chmod +x "$FAKE_CHROME_DIR/late"
check "rendered landing test gives output that Chrome writes after it exits" \
"CHROME_PATH='$FAKE_CHROME_DIR/late' node '$IDSTACK_DIR/test/test-rendered-landing.js'" \
1 "exit code 4.*fake-chrome-late-output"
rm -rf "$FAKE_CHROME_DIR"
# The fixed text that the extension shows must obey ASD-STE100: the demo audits,
# the Markdown export, the rendered findings, the error messages and the side-panel
Expand All @@ -442,6 +471,7 @@ else
echo " SKIP: responsive landing page tests (node not installed)"
echo " SKIP: rendered landing page tests (node not installed)"
echo " SKIP: rendered landing startup report (node not installed)"
echo " SKIP: rendered landing late-output report (node not installed)"
echo " SKIP: extension STE checks (node not installed)"
fi

Expand Down
88 changes: 88 additions & 0 deletions test/test-citation-tiers.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,88 @@
#!/usr/bin/env python3
"""Unit tests for test/check-citation-tiers.py.

Each test builds a small repo with an evidence/references.md and one skill
template, then runs the checker on it. The forms here are the ones a line-by-line
match missed: codes separated by commas, and a citation that a line break splits.
"""

import os
import shutil
import subprocess
import sys
import tempfile
import unittest

REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), '..'))
CHECKER = os.path.join(REPO_ROOT, 'test', 'check-citation-tiers.py')
REFS = ('- [Alpha-1] Author, A. (2020). A trial. T1\n'
'- [Beta-2] Author, B. (2021). An opinion. T5\n')


def run_checker(root):
return subprocess.run([sys.executable, CHECKER, root], stdout=subprocess.PIPE,
stderr=subprocess.PIPE, universal_newlines=True)


class CitationTierTest(unittest.TestCase):

def setUp(self):
self.root = tempfile.mkdtemp(prefix='idstack-test-tiers-')
os.makedirs(os.path.join(self.root, 'evidence'))
os.makedirs(os.path.join(self.root, 'skills', 'demo'))
with open(os.path.join(self.root, 'evidence', 'references.md'), 'w', encoding='utf-8') as f:
f.write(REFS)

def tearDown(self):
shutil.rmtree(self.root)

def check(self, text):
with open(os.path.join(self.root, 'skills', 'demo', 'SKILL.md.tmpl'), 'w', encoding='utf-8') as f:
f.write(text)
return run_checker(self.root)

def test_correct_citations_pass(self):
proc = self.check('One [Alpha-1] [T1]. Two [Beta-2] [T5].\n')
self.assertEqual(proc.returncode, 0, proc.stdout)

def test_wrong_tier_on_one_line_fails(self):
proc = self.check('Text [Beta-2] [T1].\n')
self.assertEqual(proc.returncode, 1)
self.assertIn('SKILL.md.tmpl:1: [Beta-2] is T5 in evidence/references.md, not T1', proc.stdout)

def test_each_code_in_a_comma_list_is_checked(self):
# The code before the comma is the one that a match without commas left out.
proc = self.check('Text [Beta-2], [Alpha-1] [T1].\n')
self.assertEqual(proc.returncode, 1)
self.assertIn('[Beta-2] is T5 in evidence/references.md, not T1', proc.stdout)

def test_comma_list_at_the_correct_tier_passes(self):
proc = self.check('Text [Beta-2], [Beta-2] [T5].\n')
self.assertEqual(proc.returncode, 0, proc.stdout)

def test_citation_split_by_a_line_break_is_checked(self):
proc = self.check('First line.\nText that ends with [Beta-2]\n [T1] and goes on.\n')
self.assertEqual(proc.returncode, 1)
self.assertIn('SKILL.md.tmpl:2: [Beta-2] is T5 in evidence/references.md, not T1', proc.stdout)

def test_split_citation_at_the_correct_tier_passes(self):
proc = self.check('Text that ends with [Beta-2]\n[T5] and goes on.\n')
self.assertEqual(proc.returncode, 0, proc.stdout)

def test_a_no_break_space_between_code_and_tier_is_seen(self):
proc = self.check('Text [Beta-2] [T1].\n')
self.assertEqual(proc.returncode, 1)
self.assertIn('[Beta-2] is T5 in evidence/references.md, not T1', proc.stdout)

def test_a_blank_line_ends_a_citation(self):
# A code at the end of one paragraph and a tier at the start of the next are not one citation.
proc = self.check('Text [Beta-2] [T5] ends with [Alpha-1]\n\n[T5] starts here.\n')
self.assertEqual(proc.returncode, 0, proc.stdout)

def test_the_repo_passes(self):
proc = run_checker(REPO_ROOT)
self.assertEqual(proc.returncode, 0, proc.stdout)


if __name__ == '__main__':
unittest.main()
Loading
Loading