From 13246bfe6c9f56dc050deff6eeb7961d41460748 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Tue, 29 Sep 2026 21:32:59 +0000 Subject: [PATCH 1/5] chore: add the Simple English skill from AminBlg/SimpleEnglish Copy the Simple English skill into .agents/skills/simple-english. Cursor and Codex load project skills from .agents/skills. The files come from commit 79b590fc8596523d92c26b1ea7e33236606ef069 with no changes. The project uses the MIT license, and LICENSE holds the license text. Co-authored-by: Ramon Niebla --- .agents/skills/simple-english/LICENSE | 21 + .agents/skills/simple-english/SKILL.md | 93 +++++ .../simple-english/references/rule-catalog.md | 165 ++++++++ .../references/strict-vocabulary.md | 77 ++++ .../references/system-prompt.md | 25 ++ .../simple-english/references/use-cases.md | 62 +++ .../simple-english/references/word-swaps.md | 58 +++ .../skills/simple-english/scripts/slop.tsv | 69 ++++ .../skills/simple-english/scripts/ste_lint.py | 372 ++++++++++++++++++ 9 files changed, 942 insertions(+) create mode 100644 .agents/skills/simple-english/LICENSE create mode 100644 .agents/skills/simple-english/SKILL.md create mode 100644 .agents/skills/simple-english/references/rule-catalog.md create mode 100644 .agents/skills/simple-english/references/strict-vocabulary.md create mode 100644 .agents/skills/simple-english/references/system-prompt.md create mode 100644 .agents/skills/simple-english/references/use-cases.md create mode 100644 .agents/skills/simple-english/references/word-swaps.md create mode 100644 .agents/skills/simple-english/scripts/slop.tsv create mode 100644 .agents/skills/simple-english/scripts/ste_lint.py diff --git a/.agents/skills/simple-english/LICENSE b/.agents/skills/simple-english/LICENSE new file mode 100644 index 00000000..e3dd42a3 --- /dev/null +++ b/.agents/skills/simple-english/LICENSE @@ -0,0 +1,21 @@ +MIT License + +Copyright (c) 2026 AminBlg + +Permission is hereby granted, free of charge, to any person obtaining a copy +of this software and associated documentation files (the "Software"), to deal +in the Software without restriction, including without limitation the rights +to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +copies of the Software, and to permit persons to whom the Software is +furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all +copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +SOFTWARE. diff --git a/.agents/skills/simple-english/SKILL.md b/.agents/skills/simple-english/SKILL.md new file mode 100644 index 00000000..573fb088 --- /dev/null +++ b/.agents/skills/simple-english/SKILL.md @@ -0,0 +1,93 @@ +--- +name: simple-english +description: | + Write or rewrite text in plain, layman-readable English in the spirit of + ASD-STE100 Simplified Technical English: short sentences, active voice, + simple tenses, one word one meaning, condition before command, every + technical term defined at first use, no AI slop. Default mode is Plain. + Strict mode applies full STE vocabulary compliance when the user names + STE, ASD-STE100, or compliance. Use for documentation, READMEs, runbooks, + procedures, error messages, release notes, incident reports, API guides, + and explanations for readers outside the field. Also use when the user + says "STE", "Simplified Technical English", "ASD-STE100", "plain English", + "layman's terms", "explain it simply", "no jargon", "de-slop", "make this + readable", "write for non-native readers", or asks for docs that translate + well. The same rules govern the reply: answer first, prose only. +license: MIT +compatibility: claude-code cursor codex gemini-cli opencode +metadata: + version: "2.1.0" + standard: ASD-STE100 Issue 9 (2025-01-15) +--- + +# Simple English + +Write plain English that a smart reader outside your field understands on one read. The rules come from ASD-STE100, the controlled language aerospace uses so a tired mechanic cannot misread an instruction. Two registers exist: the document you write or rewrite, and the reply you type in chat. Each has its own short rule set below. Nothing else in this file is optional. + +## The Document + +When asked to write or rewrite documentation, apply these rules to the prose: + +1. **Classify each passage.** Procedural text tells the reader what to do: imperative mood, 20 words per sentence, one instruction per sentence. Descriptive text explains: simple tenses, 25 words per sentence, one topic per paragraph, six sentences per paragraph at most. +2. **Never touch** code, identifiers, commands, flags, file paths, quoted errors, product names, or facts. When the source gives no number or cause, keep the general statement. +3. **Condition before command, with a comma.** "If the build fails, read the log." +4. **Simple tenses, active voice.** No present perfect ("has completed" → "completed"). No "-ing" verb after a comma (", making it easy" → new sentence). Name the actor: "You run the migration." +5. **Modals: can, will, must.** Never should, would, may, might, could. A required "should" becomes "must". An optional one is deleted. +6. **Complete grammar.** No contractions, keep articles, keep "that". Short sentences, not telegraph style. +7. **No semicolons and no em-dashes.** Write two sentences, or name the relation. +8. **One word, one meaning, for the whole document.** Use `make sure that` for check, verify, confirm, validate, ensure. Use `configuration` for config, settings, options. Break noun chains over three words with a preposition ("the timeout value for the connection pool"). +9. **State what the reader needs before you name the action.** Define a concept term at its first use, under ten words, one per sentence. Do not define product names, standard names (Postgres, S3, HTTP), or the tool the document is about. The same rule covers a fact, not just a word: name the host, the flag, or the prior step that a command depends on, instead of assuming the reader already has it. "Restart the service" becomes "Restart the `sync` service on the host that runs the job." +10. **State the fact, not its importance.** Delete words that carry no fact: simply, seamlessly, robust, powerful, comprehensive, leverage, crucial, "in order to", "it is worth noting". No "not just X, it is Y". No decorative triplets. No "in conclusion". +11. **Format for the eye, not for decoration.** No bold lead-ins, no bold as emphasis, no emoji, no heading over two sentences. A vertical list is for three or more parallel items or steps: colon on the lead-in, uppercase start, one instruction per item. +12. **Warnings: command or condition first, then the risk.** "Do not run this against production. The command deletes rows." + +Use American spelling. `references/word-swaps.md` maps the overused words to plain ones. For an error message, a runbook, an incident report, release notes, a commit message, or UI copy, read `references/use-cases.md` first: it names the mode and the pattern for each. + +**Before (real AI output):** + +> **Connection timeouts.** If sqlpipe hangs or fails with `dial tcp: i/o timeout`, check that the host running sqlpipe can reach the Postgres port (usually 5432) — this is often a security group or firewall rule blocking the connection. If you're connecting to a managed database (RDS, Cloud SQL, etc.), confirm the instance allows connections from sqlpipe's IP. + +**After (procedural, headed, numbered):** + +> ## Connection timeouts +> +> sqlpipe stops with `dial tcp: i/o timeout` when it cannot connect to the Postgres port (5432 by default). +> +> 1. Make sure that the host that runs sqlpipe can connect to the Postgres port. A firewall or security group usually blocks it. +> 2. If the database is managed (RDS, Cloud SQL), make sure that the instance accepts connections from the IP of sqlpipe. + +## The Reply + +Every chat reply, in every mode, follows these rules. Read them last, apply them first: + +1. Answer in prose. No headers, no bullet lists, no bold, no tables. A code block is legal when the reader must copy it. +2. The first sentence gives the answer or the result. Do not restate the question. +3. No em-dashes. Name the relation ("because", "but", "for example") or write two sentences. +4. Define a concept term in a few words the first time you use it: "idempotent (safe to run twice)". Do not define product names. +5. No contractions. No openers ("Certainly", "Great question") and no closers ("I hope this helps", "Let me know"). +6. Do not shorten quoted error text, security warnings, or confirmations before a destructive action. + +**Before:** The failure stems from control-plane leader election during pod churn — nothing to worry about! +**After:** The pods restarted and the queue lost its leader for a short time. It recovered without help. You do not have to do anything. + +## Self-Check Before You Deliver + +1. Reply: search for `—`, `**`, `#`, and a line that starts with `-`. Remove each one. +2. Document: count the words in your three longest sentences. Over 20 or 25, split. Search for `'`, `has been`, `should`, `may`, `;`, `—`, `, making`, `**`, `check`, `verify`, `config`, and any heading that covers fewer than three sentences. Fix each hit. Read each step: does it name a host, a flag, or a prior step the reader must already have? If not, add it or point to it. + +## Modes + +**Plain** is the default and is all of the above. **Strict** applies when the user names STE, ASD-STE100, or compliance: read `references/strict-vocabulary.md` before you draft the document, and say once that no tool guarantees compliance. The reply stays Plain in every mode. + +When asked to CHECK text instead of writing it, first open `references/rule-catalog.md`. Then report each violation as: rule number quoted from that file, the offending text, a compliant rewrite. Never cite a rule number from memory. When the user asked for compliance, end with one sentence: no tool can guarantee ASD-STE100 compliance, and the standard is a free download at asd-ste100.org. + +## Limits + +These rules are for facts and instructions, not marketing copy or brand writing: they delete persuasion by design. Say so, and offer them for the docs instead. + +## References + +- `references/rule-catalog.md` — the 53 rules of Issue 9 with software examples, for CHECK mode +- `references/strict-vocabulary.md` — the dictionary discipline for Strict mode +- `references/word-swaps.md` — slop-to-plain word map +- `references/use-cases.md` — mode and pattern for error messages, runbooks, incident reports, release notes, commits, agent prompts, UI copy, translation prep diff --git a/.agents/skills/simple-english/references/rule-catalog.md b/.agents/skills/simple-english/references/rule-catalog.md new file mode 100644 index 00000000..d2356fb1 --- /dev/null +++ b/.agents/skills/simple-english/references/rule-catalog.md @@ -0,0 +1,165 @@ +# Rule catalog: ASD-STE100 Issue 9 for software text + +Read this file for CHECK mode, for Strict mode, or when a rule number is in question. The core skill in `SKILL.md` names the rules that move output the most. This file holds the whole catalog. + +53 rules in 9 sections, paraphrased from ASD-STE100 Issue 9 with software examples. Rules marked (S) are Strict mode only (see `references/strict-vocabulary.md`). The official wording is in the free standard at asd-ste100.org. + +### Section 1 — Words (Rules 1.1-1.14) + +| Rule | Instruction | +|---|---| +| 1.1-1.4, 1.6 (S) | Use only approved words, as their listed part of speech, meaning, and form. | +| 1.5 | You can use domain words as technical nouns ("webhook", "commit", "endpoint"). | +| 1.7 | Do not use technical nouns as verbs. | +| 1.8 | Use the technical nouns of your project or industry. | +| 1.9 | When you pick a technical noun, pick a short and clear one. | +| 1.10 | No regional, slang, or jargon words as technical nouns. | +| 1.11 | One item, one name. Do not call it "config" here and "settings" there. | +| 1.12 | You can use domain verbs as technical verbs ("deploy", "compile", "merge"). The standard names computer verbs as legal: click, type, copy, paste, delete, save, install, download, update, and more. When a common verb does the same job, prefer it: "find" instead of "detect". | +| 1.13 | Do not use technical verbs as nouns. | +| 1.14 | Use American English spelling. | + +In Plain mode, rules 1.5, 1.8, and 1.12 make your domain vocabulary legal. The ones agents break are 1.7, 1.11, and 1.13. + +**Before:** You can webhook the event, then do a deploy. +**After:** Send the event to the webhook. Then deploy the service. + +### Section 2 — Multi-word nouns (Rules 2.1-2.2) + +| Rule | Instruction | +|---|---| +| 2.1 | Write multi-word nouns of three words or fewer. | +| 2.2 | When a technical noun needs more than three words, write it in full once, then give a short form or hyphenate the units. | + +Break long noun chains with prepositions (of, on, in, for): + +**Before:** the connection pool timeout configuration value +**After:** the timeout value for the connection pool + +### Section 3 — Verbs (Rules 3.1-3.7) + +| Rule | Instruction | +|---|---| +| 3.1 (S) | Use only the verb forms that the dictionary gives. | +| 3.2 | Use only: infinitive, imperative, simple present, simple past, simple future, past participle as adjective. | +| 3.3 | Use the past participle only as an adjective ("the cached response"). | +| 3.4 | No auxiliary verbs for complex constructions. No present perfect, no "is to be installed". | +| 3.5 | Use an "-ing" form only as a technical noun or inside one ("logging", "the mounting bracket"), never as a verb. | +| 3.6 | Active voice. In descriptive text, passive is legal only when the agent is unknown. To repair an agentless passive, use "you" (the reader) or "we" (your company): "Indexes are not used on this table" → "We do not use indexes on this table." | +| 3.7 | Describe an action with a verb, not a noun ("compress the file", not "perform compression of the file"). | + +**Approved modals: can, will, must. Banned: should, would, may, might, could.** +The modal ladder below routes each banned modal. This matters double for agent instructions, because models read "should" as optional. + +### Section 4 — Sentences (Rules 4.1-4.5) + +| Rule | Instruction | +|---|---| +| 4.1 | Write short and clear sentences. | +| 4.2 | Do not omit words or use contractions to shorten sentences. Keep articles, keep "that". | +| 4.3 | Use a vertical list for complex text: colon on the lead-in, uppercase start, a period only on full-sentence items, no mixed instructions and facts, no nesting. | +| 4.4 | Use connecting words between sentences on related topics ("Then", "As a result"). | +| 4.5 | Put an article (the, a, an) or a demonstrative adjective (this, these) before nouns where applicable. Exception: no article before a noun when an identifier follows it: "Restart pod web-7f9b2". | + +Rule 4.2 is the anti-terseness rule. Plain English is short sentences with complete grammar, not telegraph style: + +**Wrong shortening:** Ensure file exists before running. +**Plain:** Make sure that the file exists before you run the command. + +### Section 5 — Procedural writing (Rules 5.1-5.5) + +| Rule | Instruction | +|---|---| +| 5.1 | Maximum 20 words per sentence. Warnings and cautions included. | +| 5.2 | One instruction per sentence, unless two actions happen at the same time. A step can add one sentence for an immediate result or limit. | +| 5.3 | Write instructions in the imperative: "Run the migration." | +| 5.4 | Put a required condition before the command, divided by a comma: "If the build fails, read the log." | +| 5.5 | Notes give information, never instructions or limits. A limit belongs with its action. Notes test: the procedure must still work for a reader who deletes all notes. | + +**Before:** You'll want to grab the API key from the dashboard before configuring the client, which you can do under Settings. +**After:** Get the API key from the dashboard, under Settings. Then configure the client with this key. + +### Section 6 — Descriptive writing (Rules 6.1-6.6) + +| Rule | Instruction | +|---|---| +| 6.1 | Give information gradually: one new fact per sentence. | +| 6.2 | Use key words and phrases to give the text a logical structure. | +| 6.3 | Maximum 25 words per sentence. | +| 6.4 | Group related information in paragraphs. | +| 6.5 | One topic per paragraph. | +| 6.6 | Maximum six sentences per paragraph. | + +### Section 7 — Safety instructions (Rules 7.1-7.3) + +| Rule | Instruction | +|---|---| +| 7.1 | Use a word that shows the risk level ("WARNING" = injury, "CAUTION" = damage). If the two risks occur together, use "WARNING". | +| 7.2 | Start with a clear command or condition. | +| 7.3 | Then give the risk or the possible result. | + +Never bury the instruction after the explanation. The same pattern fits destructive CLI flags and irreversible migrations. + +**Before:** Note that data loss may occur in some circumstances if the destructive flag happens to be enabled when running against production. +**After:** CAUTION: Do not use the `--force` flag against production. The flag deletes rows that do not match the source. + +### Section 8 — Punctuation and word count (Rules 8.1-8.7) + +| Rule | Instruction | +|---|---| +| 8.1 | All standard punctuation is legal except the semicolon. Write two sentences instead. | +| 8.2 | Use hyphens to connect words that act as one unit. | +| 8.3 | Parentheses are legal for references, item numbers, abbreviations, plural forms, explanations, alternatives. | +| 8.4 | In a vertical list, the lead-in colon ends a sentence for word count. Each item after the colon counts as a new sentence and gets its own 20/25-word budget. | +| 8.5-8.7 | Count as one word each: text in parentheses, a hyphenated word, numbers, numbers with units, abbreviations, identifiers, quoted text, titles, labels, proper nouns. | + +Rule 8.6 matters for software text: `sqlpipe run --config sqlpipe.yaml` in backticks counts as one word. + +**Dashes** (this skill, not the standard). An em-dash (`—`) splices two statements and hides the logic between them. Name the relation ("because", "but", "for example") or write two sentences. A spaced or double hyphen between statements is the same dash. A range (`5–10`), a list marker, and a flag (`--force`) are not. + +### Section 9 — Writing practices (Rules 9.1-9.4, GR-1 to GR-8) + +| Rule | Instruction | +|---|---| +| 9.1 | When a word-for-word replacement does not work, restructure the sentence. | +| 9.2 (S) | Use each approved word correctly: approved meaning, approved part of speech. | +| 9.3 | Prefer the one-word verb over the phrasal verb ("decrease", not "go down"; "install", not "set up"). Strict mode: the phrasal verb is a violation. | +| 9.4 | Keep one consistent style and terminology through the whole document. | + +General recommendations: keep "that" (GR-1), primary verb first and the tool after "with" (GR-2: "Fetch the URL with curl"), clear pronoun referents (GR-3), "this + noun" (GR-4), inclusive language (GR-7). GR-6: "e.g." → "for example", "i.e." → "that is", delete "etc." and name the items. + +### The modal ladder + +| You wrote | Write instead | +|---|---| +| should (requirement) | must | +| should (recommendation) | Delete it, or state it as fact: "X is better because Y." | +| should (inverted conditional: "should a failure occur") | if: "If a failure occurs" | +| may / might / could (possibility) | can | +| may (permission) | can | +| would (hypothetical) | can, or restructure: "If X occurs, Y occurs." | + +## Signs of AI Writing + +AI text drifts in known directions (Wikipedia "Signs of AI writing"). The rules above remove some already. Guard against the rest by direction, in documents and replies alike: + +- Inflated significance: no "vital", "crucial", "a testament". State the fact. +- Negative parallelism: no "not just X, it is Y". +- Rule of three: no decorative triplets. +- Vague attribution: no "studies show". Name the source, or drop the claim. +- False ranges: no "ranging from X to Y" without real limits. +- Restating summaries: no "in conclusion" paragraphs. +- Editorializing asides: no "it is important to note". +- Collaborative leftovers: no "I hope this helps", no "Let me know". +- Formatting habits: no bold as decoration, no bold lead-ins, no emoji as structure, no heading for two sentences. + +For the specific overused words, `references/word-swaps.md` maps each one to a plain replacement. If a word carries no fact, delete it instead. + +## Word Choice + +One word, one meaning, one part of speech, for the whole document (Rules 1.11, 9.4). + +- The settings file is `configuration`, never config, settings, or options in the same document. +- The verify concept is `make sure that`, never check, verify, confirm, validate, or ensure as verbs. Strict mode routes the rest with `references/strict-vocabulary.md`. +- Common swaps: however → but, therefore → as a result, since (= because) → because, perform → do, avoid → prevent, repeat → do again, acceptable → permitted, now → delete it. + diff --git a/.agents/skills/simple-english/references/strict-vocabulary.md b/.agents/skills/simple-english/references/strict-vocabulary.md new file mode 100644 index 00000000..2233a6da --- /dev/null +++ b/.agents/skills/simple-english/references/strict-vocabulary.md @@ -0,0 +1,77 @@ +# Strict mode: the dictionary discipline + +Read this file when the user names STE, ASD-STE100, or compliance. Strict mode adds the dictionary rules to the document. It does not change the reply to the user, which stays in Plain mode. + +The official dictionary (about 900 approved words and 1,200 rejected words with their alternatives) is copyrighted by ASD and is not reproduced here. This file gives the rules that depend on it, the rulings that software writers meet most, and the words the standard names as recurring errors. Tell the user, one time per conversation and in one sentence, that this index is lossy and that full compliance needs the official dictionary (free at asd-ste100.org). + +## Rules that exist only with the dictionary + +| Rule | Instruction | +|---|---| +| 1.1 | Use only approved words, technical nouns, or technical verbs. | +| 1.2 | Use an approved word only as its listed part of speech. | +| 1.3 | Use an approved word only with its approved meaning. | +| 1.4 | Use only the approved forms of verbs and adjectives. | +| 1.6 | Use an unapproved word only when it is a technical noun or part of one. | +| 3.1 | Use only the verb forms that the dictionary gives. | +| 9.2 | Use each approved word correctly: approved meaning, approved part of speech. | +| GR-5 | Avoid false friends: words that look like a word in the reader's language but mean something else. | +| GR-8 | Use the possessive apostrophe only when you are sure that it is correct. If unsure, write "the file of the user". | + +Issue 9 adds a quick-reference list of approved verbs in the dictionary introduction. Check your verbs against it. + +## Part-of-speech rulings + +| Word | Ruling | +|---|---| +| test, check, work | Noun only. "Do a test", not "test the pump". "Check that X" becomes "make sure that X". | +| oil | Technical noun only. For the verb, the dictionary gives "lubricate": "Lubricate the linkage with oil." | +| help | Verb only. For the noun, the dictionary gives "aid": "with the aid of". | +| fall (noun) | Rejected. Use "decrease" for a reduction in value. Use "fall" (verb) only for movement downward by gravity: "Make sure that the tools do not fall into the engine." | +| follow | "To come after" only, never "obey". Write "obey the instructions". | +| above, below | Physical positions only. For limits write "more than", "less than". | + +## Dictionary rulings on common software verbs + +The standard has already chosen. Use the approved word. + +| You wrote | Dictionary status | Use instead | +|---|---|---| +| check (verb), verify, confirm, ensure | All rejected as verbs | Route by intent: `make sure that` (a state), `examine` (look for faults: "examine the log"), `measure` (get a value), or the noun: "do a check of". | +| validate | Not in the dictionary | Legal as a technical verb (Rule 1.12), or replace with `make sure that`. | +| delete, drop (verb), destroy | All rejected as dictionary verbs | `erase` (data), `remove` (physical). In computer contexts `delete` is also a legal technical verb (Rule 1.12). Do not use `drop` or `destroy`. | +| remove | Approved verb | Keep it. | +| run, execute | Both rejected | `operate` for run, `do` for execute. | +| invoke, launch | Not in the dictionary | Legal as technical verbs (Rule 1.12). | +| display (verb), render, present (verb) | All rejected | `show` covers most software cases. Official alternatives: display → `show`, render → `make`, present → `give` or `show`. | +| issue | Not in the dictionary | Use as a technical noun, or replace with `problem` (approved). | +| failure | Rejected in general use; approved as a technical noun for performance loss | Use only for a performance error: "a failure of the pump". | +| error, problem | Approved nouns | Keep them. | + +## Recurring errors the standard names + +The dictionary introduction lists the words that writers get wrong most often. This is the software-relevant set, as rulings only. + +| You wrote | STE writes | +|---|---| +| however | but | +| therefore | thus, as a result | +| since (= because) | because | +| any | Delete it, or restructure: "if you have any questions" → "if you have questions" | +| now | at this time. Better, delete it: "now start the service" → "start the service" | +| need to, have to | Imperative in procedures ("install"); "it is necessary to" in descriptive text | +| perform | do | +| insert | put (but SQL `INSERT` stays: it is quoted text) | +| reach | get, get to | +| avoid | prevent | +| repeat | do … again | +| acceptable | permitted. Better, give the limit: "a latency of less than 200 ms" | +| complete (adjective) | completed | +| the example below, the section above | Name the target, or put the reference after it: "the example that follows" | + +## Strict self-check + +Add these two steps to the self-check in SKILL.md: + +1. Search the draft for every verb in the two tables above. Replace each hit with the approved word. +2. Search for the phrasal verbs you built ("set up", "go down"). Replace each with the one-word verb (Rule 9.3: "install", "decrease"). diff --git a/.agents/skills/simple-english/references/system-prompt.md b/.agents/skills/simple-english/references/system-prompt.md new file mode 100644 index 00000000..d01747ee --- /dev/null +++ b/.agents/skills/simple-english/references/system-prompt.md @@ -0,0 +1,25 @@ +# Standalone system prompt + +For harnesses without SKILL.md support: paste this block into your system prompt, custom instructions, AGENTS.md, or `.cursorrules`. It is the condensed version of the full skill. + +--- + +Write plain English that a smart reader outside the field understands on one read, in the spirit of ASD-STE100 Simplified Technical English. Two registers, each with its own rules. + +THE DOCUMENT (documentation, READMEs, runbooks, error messages, release notes, reports, commit messages). Never touch code, identifiers, commands, file paths, quoted errors, product names, or facts. Classify each passage. Procedural text tells the reader what to do: imperative mood, 20 words per sentence, one instruction per sentence. Descriptive text explains: simple tenses, 25 words per sentence, one topic per paragraph, six sentences per paragraph at most. Condition before command, with a comma: "If the build fails, read the log." Simple tenses, active voice: no present perfect ("has completed" → "completed"), no "-ing" verb after a comma. Name the actor: "You run the migration." Modals: can, will, must. Never should, would, may, might, could. Complete grammar: no contractions, keep articles, keep "that". No semicolons and no em-dashes. One word, one meaning: `make sure that` for check, verify, confirm, validate, ensure. `configuration` for config, settings, options. Noun chains of three words at most. Define a concept term at its first use, under ten words, one per sentence. Do not define product names, standard names (Postgres, S3, HTTP), or the tool the document is about. The same rule covers a fact, not just a word: name the host, the flag, or the prior step that a command depends on, instead of assuming the reader already has it. State the fact, not its importance: delete simply, seamlessly, robust, powerful, comprehensive, leverage, crucial, "in order to", "it is worth noting". No "not just X, it is Y", no decorative triplets, no "in conclusion". No bold lead-ins, no bold as emphasis, no emoji, no heading over two sentences. A vertical list is for three or more parallel items or steps. Warnings: command or condition first, then the risk. American spelling. + +SELF-CHECK. Document: count the words in your three longest sentences, split any over the limit. Search for "'", "has been", "should", "may", ";", "—", ", making", "check", "verify", "config". + +STRICT MODE. If the user names STE, ASD-STE100, or compliance, also apply the STE dictionary to the document: "make sure that" for check/verify/confirm, "operate" for run, "do" for execute, "show" for display, "but" for however, "because" for since. Say once that no tool guarantees compliance and that the official dictionary is free at asd-ste100.org. + +Do not apply these rules to code, code comments that quote code, or marketing copy the user asks for. + +THE REPLY (every chat reply, in every mode). Answer in prose: no headers, no bullet lists, no bold, no tables. A code block is legal when the reader must copy it. The first sentence gives the answer or the result. Do not restate the question. No em-dashes: name the relation ("because", "but", "for example") or write two sentences. Define a concept term in a few words the first time ("idempotent (safe to run twice)"), never a product name. No contractions. No openers ("Certainly", "Great question") and no closers ("I hope this helps", "Let me know"). Do not shorten quoted error text, security warnings, or confirmations before a destructive action. + +--- + +## Word-budget version (~60 tokens) + +For tight system prompts: + +> Replies: prose only, five sentences max, answer first, no headers, bullets, bold, tables, or em-dashes, define terms, no contractions. Documents: ASD-STE100 style, 20 words per instruction sentence, 25 per description, imperative steps, condition before command, simple tenses, active voice, no should/would/may/might, one word per meaning, no semicolons or em-dashes, no filler, code exact. diff --git a/.agents/skills/simple-english/references/use-cases.md b/.agents/skills/simple-english/references/use-cases.md new file mode 100644 index 00000000..a98c215a --- /dev/null +++ b/.agents/skills/simple-english/references/use-cases.md @@ -0,0 +1,62 @@ +# Use cases beyond documentation + +STE was built for aircraft maintenance manuals. The same properties transfer to any text where a misreading has a cost: one meaning per word, short sentences, condition-first commands. Each case below names the mode and the adaptations. + +## Error messages and CLI output + +Mode: procedural. An error message is an instruction to a stressed reader at 2 a.m., so it is the highest-value target. + +Pattern: state what happened (simple past), state the cause if known, give the command or condition that fixes it. + +> Before: Oops! Something went wrong while attempting to establish a connection. Please ensure your credentials are properly configured and try again. +> After: Connection to the database failed. The password for user `app` was not correct. Set `DB_PASSWORD` and connect again. + +## Runbooks and standard operating procedures + +Mode: procedural, with the 20-word limit enforced hard. An on-call runbook is a maintenance manual, which is what STE was made for. + +- Every step is imperative, one instruction per step, condition first. +- A warning comes before its step: command first, risk second. +- An operator under pager stress reads each sentence once, so the 20-word limit is not negotiable. + +## Incident reports and postmortems + +Mode: descriptive, simple past only. A timeline in present perfect ("we have identified") hides when things happened. + +> Before: We have identified an issue that may have impacted some users' ability to access the service. +> After: Between 14:02 and 14:31 UTC, 12% of requests failed. A deploy at 14:00 removed the cache warmup step. + +STE bans hedges such as "may have impacted". The report states what is known and says "unknown" for the rest. It reads more honest because it is. + +## Commit messages and PR descriptions + +Mode: imperative subject line, descriptive body. The convention already matches STE. Apply the word swaps and the 25-word limit to the body. Delete "this PR aims to". + +## API changelogs and release notes + +Mode: descriptive. One entry, one change, one sentence where possible. A "Breaking:" entry follows the warning pattern, command first: "Update your calls to `v2/users`. The `name` field split into `first_name` and `last_name`." + +## Instructions for AI agents (prompts, AGENTS.md, skills) + +Mode: procedural. A system prompt is a procedure for a reader that cannot ask questions, which is the exact reader STE was designed for. + +- One instruction per sentence keeps each rule quotable and hard to half-follow. +- One word, one meaning stops the model from treating "check", "verify", and "validate" as three operations. +- A condition first ("If the build fails, stop") beats a trailing condition, which models drop. +- No "should". A model reads "should" as optional. Write "must" or delete the rule. + +## Support macros and status-page updates + +Mode: descriptive, 25-word limit. Non-native readers are the majority of many user bases. Not "we sincerely apologize for any inconvenience this may have caused" but "The API was down for 18 minutes. Uploads made during this time were saved and will process today." + +## Translation and localization prep + +Mode: strict. The original purpose of STE was English that non-native maintenance crews can read, and it doubles as pre-editing for machine translation. One meaning per word plus complete grammar (articles, "that") removes most translation ambiguity. If your docs get localized, STE cuts the error rate and the cost. + +## UI copy and empty states + +Mode: procedural, hard length limits. Buttons and labels are technical names and are exempt. Body copy follows the rules: "No projects yet. Create a project to start." + +## Where STE does not fit + +Marketing pages, launch posts, blog voice, brand writing. STE deletes persuasion on purpose. Write those in your own voice. Then use STE for the docs that the landing page links to. diff --git a/.agents/skills/simple-english/references/word-swaps.md b/.agents/skills/simple-english/references/word-swaps.md new file mode 100644 index 00000000..c92286d7 --- /dev/null +++ b/.agents/skills/simple-english/references/word-swaps.md @@ -0,0 +1,58 @@ +# Slop-to-simple substitutions + +This table is ours, not the ASD dictionary. It maps the words AI-generated docs overuse to plain replacements. If the word carries no fact, delete it instead of replacing it. + +| Slop | Write instead | +|---|---| +| leverage, utilize | use | +| in order to | to | +| prior to | before | +| ensure | make sure that | +| it is worth noting that | (delete) | +| it's important to | (delete — state the fact) | +| simply, just, easily, seamless, seamlessly, effortlessly | (delete) | +| robust, powerful, comprehensive, performant | (delete, or give the measurable property) | +| functionality | function, feature | +| enables you to, allows you to | you can | +| is designed to, aims to | (delete — say what it does) | +| facilitate | help, make possible | +| dive into, delve into | read, examine | +| when it comes to | for | +| in the event that | if | +| due to the fact that | because | +| as needed, as necessary | (state the condition) | +| and/or | Pick one, or write "X, or Y, or both" | +| e.g. / i.e. / etc. | for example / that is / (name the items) | +| gracefully handles | (say what it does: "retries three times, then stops") | +| out of the box | by default | +| under the hood | internally | +| blazingly fast | fast (give the number) / (delete) | +| streamline | make simpler, make faster | +| plethora, myriad | many | +| addresses the issue, tackles | corrects the fault, removes the error | +| pivotal, crucial, crucially, paramount | important | +| tapestry, testament, synergy | (delete) | +| interplay | interaction (or delete) | +| intricate | complex | +| vibrant, nuanced, multifaceted | (delete, or name the parts) | +| realm, landscape (metaphorical) | area | +| groundbreaking, cutting-edge, state-of-the-art, innovative, unprecedented | new (or delete) | +| transformative, game-changer | (delete — say what changes) | +| revolutionize | change | +| showcase, underscore, emphasize | show | +| foster, empower, bolster | help, support, let | +| harness | use | +| enhance | improve | +| elevate | increase | +| furthermore, moreover | also | +| in conclusion, in summary, at the end of the day | (delete) | +| embark, endeavor | start, try | +| meticulous, meticulously | careful, carefully | +| holistic | full | +| paradigm | model | +| navigate (metaphorical) | go to | +| boasts | has | +| nestled, in the heart of | (delete — give the location or the fact) | +| bustling | busy | +| that being said, notwithstanding | but | +| I hope this helps, let's dive in | (delete) | diff --git a/.agents/skills/simple-english/scripts/slop.tsv b/.agents/skills/simple-english/scripts/slop.tsv new file mode 100644 index 00000000..cce8ecbb --- /dev/null +++ b/.agents/skills/simple-english/scripts/slop.tsv @@ -0,0 +1,69 @@ +delve 28 examine +pivotal 27 important +tapestry 26 (delete) +vibrant 25 (delete) +robust 22 (give the measurable property) +intricate 21 complex +crucial 21 important +realm 20 area +groundbreaking 20 new +seamless 18 (delete) +transformative 17 (delete) +testament 17 (delete) +leverage 17 use +cutting-edge 17 new +in conclusion 17 (delete) +showcase 16 show +foster 16 help +furthermore 16 also +underscore 15 show +synergy 15 (delete) +interplay 15 interaction +harness 15 use +enhance 15 improve +elevate 15 increase +moreover 15 also +streamline 14 make simpler +profound 14 large +multifaceted 14 (name the parts) +landscape 14 area +game-changer 14 (delete) +unleash 13 release +paramount 13 most important +nuanced 13 (delete) +meticulous 13 careful +in order to 13 to +holistic 13 full +revolutionize 12 change +empower 12 let +embark 12 start +consequently 12 as a result +unprecedented 11 new +paradigm 11 model +innovative 11 new +enduring 11 permanent +valuable 10 useful +unparalleled 10 (delete) +remarkable 10 (delete) +emphasize 10 show +commendable 10 good +bustling 10 busy +at the end of the day 10 (delete) +in summary 10 (delete) +due to the fact that 10 because +state-of-the-art 9 new +renowned 9 known +nestled 9 in +endeavor 9 try +elucidate 9 explain +bolster 9 support +let's dive in 9 (delete) +notably 8 (delete) +navigate 8 go to +load-bearing 8 important +in the heart of 8 in +ever-evolving 8 (delete) +boast 8 have +that being said 8 but +notwithstanding 8 but +i hope this helps 8 (delete) diff --git a/.agents/skills/simple-english/scripts/ste_lint.py b/.agents/skills/simple-english/scripts/ste_lint.py new file mode 100644 index 00000000..50c3f76b --- /dev/null +++ b/.agents/skills/simple-english/scripts/ste_lint.py @@ -0,0 +1,372 @@ +#!/usr/bin/env python3 +"""Deterministic ASD-STE100 violation counter for benchmark runs. + +Counts mechanical violations that a regex can catch: sentence length, +contractions, banned modals, perfect tenses, "-ing" clauses, semicolons, +em-dashes, Latin abbreviations, slop words, trailing conditions, synonym +rotation. + +Known ceiling: this is a regex pass, not a grammar parser. It undercounts +(no passive-voice detection, no part-of-speech checks) and it can miscount +sentence bounds in unusual markdown. Numbers from this tool are comparable +between two texts run through the same version; they are not a compliance +verdict. No tool can guarantee STE compliance. + +Usage: + python3 ste_lint.py --type procedural file.md + cat text.md | python3 ste_lint.py --type descriptive - + python3 ste_lint.py --self-test +""" +import json +import pathlib +import re +import sys +from collections import Counter + +BANNED_MODALS = re.compile(r"\b(should|would|may|might|could)\b", re.I) +PERFECT = re.compile(r"\b(has|have|had)\s+been\b|\b(has|have)\s+\w+ed\b", re.I) +CONTRACTION = re.compile(r"\b\w+(n't|'ll|'re|'ve|'d)\b|\bit's\b|\byou're\b", re.I) +ING_CLAUSE = re.compile(r",\s*(mak|allow|enabl|ensur|highlight|creat|provid|offer|help|reduc|improv|lead|caus|result)ing\b", re.I) +LATIN = re.compile(r"\b(e\.g\.|i\.e\.|etc\.?)(?=[\s,)]|$)", re.I) +SLOP_CORE = re.compile( + r"\b(simply|seamlessly|effortlessly|robust|leverag\w*|utiliz\w*|" + r"comprehensive|powerful|blazingly|streamlin\w*|facilitat\w*|" + r"performant|plethora|myriad|delve|crucial|pivotal)\b", re.I) +SLOP_TSV = pathlib.Path(__file__).resolve().parent / "slop.tsv" + + +def slop_pattern(): + """Union of the measured core list and evals/slop.tsv (term, count, swap). + + The TSV is the 69-term LLM-tell lexicon: words named by 8 or more of 122 + published ban lists. Falls back to the core list when the file is absent. + """ + terms = [] + if SLOP_TSV.exists(): + for line in SLOP_TSV.read_text(encoding="utf-8").splitlines(): + term = line.split("\t")[0].strip().lower() + if term: + terms.append(re.escape(term).replace(r"\ ", r"\s+") + r"\w*") + if not terms: + return SLOP_CORE + return re.compile(SLOP_CORE.pattern[:-len(r")\b")] + "|" + "|".join(terms) + r")\b", re.I) + + +SLOP = slop_pattern() +# Linear scan; lint() checks ">= 4 chars before the match" instead of the old +# prefix pattern, whose backtracking was quadratic on long sentences +# (a punctuation-free 8,000-word input took ~7s; now sub-millisecond). +TRAILING_COND = re.compile(r"\s(if|when)\s", re.I) +DASH = re.compile(r"—|(?= 2] + + +def lint_detail(text, text_type): + """Every hit lint() counts, with its matched text and line number. + + strip_code() keeps every original newline (it blanks or rewrites text in + place, never deletes a line), so a line number counted in the stripped + body is the same line number in the caller's original text. Locating a + whole sentence uses its start offset in body, found once per sentence + with str.find(), which is safe here because sentences() never returns + the same sentence text twice for two different source positions in a + single lint pass (each split fragment keeps its surrounding words). + """ + body = strip_code(text) + limit = LIMITS[text_type] + hits = [] + + def add(category, m, snippet=None): + line = body.count("\n", 0, m.start()) + 1 + hits.append({"category": category, "text": (snippet or m.group(0)).strip(), "line": line}) + + def locate(sentence, search_from): + """The sentence's start offset in body, at or after search_from.""" + start = body.find(sentence, search_from) + return start if start != -1 else search_from + + def add_sentence(category, sentence, start): + line = body.count("\n", 0, start) + 1 + text_out = sentence if len(sentence) <= 80 else sentence[:80] + "…" + hits.append({"category": category, "text": text_out, "line": line}) + + pos = 0 + for s in sentences(body): + pos = locate(s, pos) + n = len(s.split()) + if n > limit: + add_sentence("sentence_over_limit", s, pos) + m = TRAILING_COND.search(s) + if m: + line_start = s.rfind("\n", 0, m.start()) + 1 + if m.start() - line_start >= 4 and not re.match(r"^(if|when)\b", s, re.I): + add_sentence("trailing_condition", s, pos) + pos += max(len(s), 1) + + for m in CONTRACTION.finditer(body): + add("contraction", m) + for m in BANNED_MODALS.finditer(body): + add("banned_modal", m) + for m in PERFECT.finditer(body): + add("perfect_tense", m) + for m in ING_CLAUSE.finditer(body): + add("ing_clause", m) + for m in re.finditer(";", body): + add("semicolon", m) + for m in DASH.finditer(body): + add("em_dash", m) + for m in LATIN.finditer(body): + add("latin_abbrev", m) + for m in SLOP.finditer(body): + add("slop_word", m) + for name, rx in ROTATION_SETS: + seen = {} + for m in rx.finditer(body): + stem = m.group(1).lower().rstrip("s") + seen.setdefault(stem, m) + for m in list(seen.values())[1:]: + add("synonym_rotation", m, f"{m.group(0)} ({name})") + + return sorted(hits, key=lambda h: h["line"]) + + +def lint(text, text_type): + body = strip_code(text) + sents = sentences(body) + limit = LIMITS[text_type] + counts = {} + lengths = [len(s.split()) for s in sents] + counts["sentence_over_limit"] = sum(1 for n in lengths if n > limit) + counts["contraction"] = len(CONTRACTION.findall(body)) + counts["banned_modal"] = len(BANNED_MODALS.findall(body)) + counts["perfect_tense"] = len([m for m in PERFECT.finditer(body)]) + counts["ing_clause"] = len(ING_CLAUSE.findall(body)) + counts["semicolon"] = body.count(";") + counts["em_dash"] = len(DASH.findall(body)) + counts["latin_abbrev"] = len(LATIN.findall(body)) + counts["slop_word"] = len(SLOP.findall(body)) + def trailing_cond(s): + m = TRAILING_COND.search(s) + if not m: + return False + # The whitespace before "if" may be a newline (a wrapped sentence), but + # the 4-char prefix must sit on the same line as that whitespace. A + # heading, a blank line, then "If ..." is condition-first, not trailing. + line_start = s.rfind("\n", 0, m.start()) + 1 + return m.start() - line_start >= 4 and not re.match(r"^(if|when)\b", s, re.I) + + counts["trailing_condition"] = sum(1 for s in sents if trailing_cond(s)) + rotation = 0 + for _, rx in ROTATION_SETS: + stems = {m.group(1).lower().rstrip("s") for m in rx.finditer(body)} + if len(stems) > 1: + rotation += len(stems) - 1 + counts["synonym_rotation"] = rotation + words = max(1, len(body.split())) + total = sum(counts.values()) + return { + "type": text_type, + "words": words, + "sentences": len(sents), + "mean_sentence_words": round(sum(lengths) / max(1, len(lengths)), 1), + "longest_sentence_words": max(lengths, default=0), + "violations": counts, + "violations_total": total, + "violations_per_100w": round(100.0 * total / words, 2), + } + + +SLOP_FIXTURE = """Leveraging our robust retry mechanism, failed uploads are automatically +reattempted, ensuring data integrity is maintained throughout the entire process which has +been designed from the ground up to gracefully handle even the most challenging network +interruptions. You should verify your credentials; it's also worth checking the settings, +e.g. the timeout config. Contact support if the problem persists.""" + +CLEAN_FIXTURE = """The system retries a failed upload automatically. This process keeps the data correct. + +If failures continue, make sure that your credentials are correct. If the problem continues, contact support.""" + +TABLE_FIXTURE = """\ +| Column A | Column B | Column C | +|----------|----------|----------| +| Cell one that is very long and has many words | Cell two that is also quite long | Cell three | +| More data here in this cell | And more data here too | And even more here | +""" + +LIST_FIXTURE = """\ +The following items are available: + +- First item without a period at the end of the line +- Second item without a period at the end of the line +- Third item without a period at the end of the line + +Following prose sentence. +""" + +LABEL_LIST_FIXTURE = """\ +**Helix owns:** + +- Learner authentication and session management +- Identity verification and multi-factor authentication +- Privacy controls and consent management +""" + +# Only the first three dashes must be flagged as logic junctions. +DASH_FIXTURE = """The deploy failed — the disk was full. +The upload failed -- the token expired. +The retry failed - the port was closed. +Do not use --force against production. +The window is 5 - 10 minutes. +The range is 5–10 minutes, over the 2024–2025 season. +Write x - y = z on the board. +Use the `--config sqlpipe.yaml` flag. +Remove the panel: + - Loosen the four bolts. +""" + + +BOLD = re.compile(r"\*\*[^*\n]+\*\*") +HEADER = re.compile(r"^#{1,6}\s", re.M) +BULLET = re.compile(r"^\s*([-*+]|\d+[.)])\s", re.M) + + +def reader_check(text): + """What a reader sees in a chat reply. + + `sentences` counts list items and table rows as sentences, but it is a report, not a + limit: the reply register sets no sentence cap, so it is not part of `visible_total`. + """ + text = text.replace("\r\n", "\n") + prose = re.sub(r"```.*?```", " ", text, flags=re.S) + prose = re.sub(r"`[^`\n]+`", " CODESPAN ", prose) + prose_no_md = re.sub(r"^\s*(#{1,6}\s|[-*+]\s|\d+[.)]\s|\|)", "", prose, flags=re.M) + prose_no_md = re.sub(r"^\s*[\s:|-]+$", "", prose_no_md, flags=re.M) # table separator rows + sents = [p for p in re.split(r"(?<=[.!?])[\"')\]]*\s+|\n+", prose_no_md) if len(p.strip().split()) >= 2] + counts = { + "sentences": len(sents), + "em_dash": len(DASH.findall(prose)), + "bold_spans": len(BOLD.findall(prose)), + "headers": len(HEADER.findall(prose)), + "bullets": len(BULLET.findall(prose)), + "contraction": len(CONTRACTION.findall(prose)), + } + words = max(1, len(prose_no_md.split())) + visible = counts["em_dash"] + counts["bold_spans"] + counts["headers"] + counts["bullets"] + return {"type": "reply", "words": words, "counts": counts, "visible_total": visible} + + +REPLY_FIXTURE = """**Yes** — it is bad. + +## Why +- Lag grows. +- Users wait. + +Check it now. Then scale.""" + + +def self_test(): + slop = lint(SLOP_FIXTURE, "procedural") + clean = lint(CLEAN_FIXTURE, "procedural") + dashes = lint(DASH_FIXTURE, "procedural") + assert slop["violations"]["sentence_over_limit"] >= 1, slop + assert slop["violations"]["banned_modal"] >= 1, slop + assert slop["violations"]["contraction"] >= 1, slop + assert slop["violations"]["perfect_tense"] >= 1, slop + assert slop["violations"]["ing_clause"] >= 1, slop + assert slop["violations"]["semicolon"] == 1, slop + assert slop["violations"]["latin_abbrev"] >= 1, slop + assert slop["violations"]["slop_word"] >= 2, slop + assert slop["violations"]["trailing_condition"] >= 1, slop + assert slop["violations"]["synonym_rotation"] >= 1, slop + assert clean["violations_total"] == 0, clean + assert lint("We delve into the landscape.", "descriptive")["violations"]["slop_word"] == 2 + assert dashes["violations"]["em_dash"] == 3, dashes + r = reader_check(REPLY_FIXTURE)["counts"] + assert (r["sentences"], r["em_dash"], r["bold_spans"], r["headers"], r["bullets"]) == (5, 1, 1, 1, 2), r + assert reader_check("Yes. It is bad because lag grows. Scale now.")["visible_total"] == 0 + tricky = "He said \"stop now.\" Then (it broke.) All done.\n\n```\n# not a header\n**not bold**\n- not a bullet\n```\n" + t = reader_check(tricky)["counts"] + assert (t["sentences"], t["headers"], t["bold_spans"], t["bullets"]) == (3, 0, 0, 0), t + cell = lint("| Column | Value |\n|---|---|\n| You should restart it | ok |\n", "descriptive")["violations"] + assert cell["banned_modal"] == 1 and cell["sentence_over_limit"] == 0, cell + assert sentences("- Loosen the bolts.\n- Then go on") == ["Loosen the bolts.", "Then go on."] + assert "visible_total" in reader_check("Yes.") and "violations_total" in lint("Yes.", "descriptive") + # Table rows and list items must not produce false sentence_over_limit hits. + assert lint(TABLE_FIXTURE, "descriptive")["violations"]["sentence_over_limit"] == 0, \ + "table rows produced false sentence_over_limit" + assert lint(LIST_FIXTURE, "descriptive")["violations"]["sentence_over_limit"] == 0, \ + "bullet list produced false sentence_over_limit" + assert lint(LABEL_LIST_FIXTURE, "descriptive")["violations"]["sentence_over_limit"] == 0, \ + "bold label + list produced false sentence_over_limit" + # A genuine long prose sentence must still flag. + long_prose = "a " * 30 # 30 repetitions of the same word, one space each + assert lint(long_prose, "descriptive")["violations"]["sentence_over_limit"] >= 1, \ + "genuine long sentence was not flagged" + detail = lint_detail(SLOP_FIXTURE, "procedural") + assert len(detail) == slop["violations_total"], (len(detail), slop["violations_total"]) + assert all(h["line"] >= 1 and h["text"] for h in detail), detail + detail_counts = Counter(h["category"] for h in detail) + assert detail_counts == {k: v for k, v in slop["violations"].items() if v}, (detail_counts, slop["violations"]) + print("self-test OK:", slop["violations_total"], "violations in slop fixture, 0 in clean") + + +USAGE = "usage: ste_lint.py [--type procedural|descriptive|reply] [--gate] (FILE|-) | --self-test" + + +def main(): + args = sys.argv[1:] + if "--self-test" in args: + self_test() + return 0 + gate = "--gate" in args + if gate: + args.remove("--gate") + text_type = "descriptive" + if "--type" in args: + i = args.index("--type") + if i + 1 >= len(args): + sys.exit("missing value after --type\n" + USAGE) + text_type = args[i + 1] + del args[i:i + 2] + if text_type != "reply" and text_type not in LIMITS: + sys.exit("unknown --type %r (expected procedural or descriptive)\n%s" % (text_type, USAGE)) + if len(args) != 1: + sys.exit(USAGE) + src = args[0] + if src == "-": + text = sys.stdin.read() + else: + try: + with open(src, encoding="utf-8") as fh: + text = fh.read() + except OSError as err: + sys.exit(str(err)) + report = reader_check(text) if text_type == "reply" else lint(text, text_type) + print(json.dumps(report, indent=2)) + total = report["visible_total"] if text_type == "reply" else report["violations_total"] + return 1 if gate and total else 0 + + +if __name__ == "__main__": + sys.exit(main()) From 06204725855514b6f465fbc807c4a9a0b67c24e8 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Tue, 29 Sep 2026 21:32:59 +0000 Subject: [PATCH 2/5] fix: keep line numbers correct in the Simple English linter The upstream strip_code function replaced each code block and some table rows with one space. It removed their newlines, so each line number after a code block or a table was too small. In this repository, 34 hits went to the wrong line. strip_code now keeps the newlines that it removes. The violation counts do not change. Co-authored-by: Ramon Niebla --- .agents/skills/simple-english/scripts/ste_lint.py | 12 +++++++++--- 1 file changed, 9 insertions(+), 3 deletions(-) diff --git a/.agents/skills/simple-english/scripts/ste_lint.py b/.agents/skills/simple-english/scripts/ste_lint.py index 50c3f76b..5b099c01 100644 --- a/.agents/skills/simple-english/scripts/ste_lint.py +++ b/.agents/skills/simple-english/scripts/ste_lint.py @@ -65,13 +65,19 @@ def slop_pattern(): LIMITS = {"procedural": 20, "descriptive": 25} +def _newlines(text): + return "\n" * text.count("\n") + + def strip_code(text): - text = re.sub(r"```.*?```", " ", text, flags=re.S) + # ldcli change: each replacement keeps the newlines that it removes, so the line + # numbers from lint_detail() match the original text after code blocks and tables. + text = re.sub(r"```.*?```", lambda m: " " + _newlines(m.group(0)), text, flags=re.S) text = re.sub(r"`[^`\n]+`", " CODESPAN ", text) # one word per Rule 8.6 text = re.sub(r"^#+\s.*$", " ", text, flags=re.M) # headings exempt (titles, 8.6) text = re.sub(r"https?://\S+", " URL ", text) - text = re.sub(r"^\s*\|[\s:|-]+\|\s*$", " ", text, flags=re.M) # table separator rows - text = re.sub(r"^\s*\|(.*)\|\s*$", lambda m: ". ".join(c.strip() for c in m.group(1).split("|") if c.strip()) + ". ", text, flags=re.M) # each cell is its own unit, still linted + text = re.sub(r"^(\s*)\|[\s:|-]+\|\s*$", lambda m: _newlines(m.group(1)) + " " + _newlines(m.group(0)[len(m.group(1)):]), text, flags=re.M) # table separator rows + text = re.sub(r"^(\s*)\|(.*)\|(\s*)$", lambda m: _newlines(m.group(1)) + ". ".join(c.strip() for c in m.group(2).split("|") if c.strip()) + ". " + _newlines(m.group(3)), text, flags=re.M) # each cell is its own unit, still linted return text From 1be871e507d3eed29f2bc42e82d80b8181a7ab3b Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Tue, 29 Sep 2026 21:33:08 +0000 Subject: [PATCH 3/5] feat: add Simple English hooks for Cursor, Claude Code, and Codex Add hooks that load the Simple English rules at session start and lint the Markdown files that an agent writes. The hooks are advisory, so they never block an action. One script, scripts/hook.py, serves the three tools. The lint reports only the lines that differ from the last commit, so an agent does not rewrite old text. Cursor also runs the hooks in .claude/settings.json. The script finds this case and runs only the Cursor hooks. Co-authored-by: Ramon Niebla --- .agents/skills/simple-english/UPSTREAM.md | 33 +++ .agents/skills/simple-english/scripts/hook.py | 255 ++++++++++++++++++ .../simple-english/scripts/test_hook.py | 179 ++++++++++++ .claude/settings.json | 38 +++ .codex/hooks.json | 19 ++ .cursor/hooks.json | 18 ++ 6 files changed, 542 insertions(+) create mode 100644 .agents/skills/simple-english/UPSTREAM.md create mode 100644 .agents/skills/simple-english/scripts/hook.py create mode 100644 .agents/skills/simple-english/scripts/test_hook.py create mode 100644 .claude/settings.json create mode 100644 .codex/hooks.json create mode 100644 .cursor/hooks.json diff --git a/.agents/skills/simple-english/UPSTREAM.md b/.agents/skills/simple-english/UPSTREAM.md new file mode 100644 index 00000000..9bac87ce --- /dev/null +++ b/.agents/skills/simple-english/UPSTREAM.md @@ -0,0 +1,33 @@ +# Source of this skill + +This folder holds a copy of the Simple English skill from [AminBlg/SimpleEnglish](https://github.com/AminBlg/SimpleEnglish), at commit `79b590fc8596523d92c26b1ea7e33236606ef069`. The project uses the MIT license. The license text is in `LICENSE`. + +## Files from the upstream project + +These files are copies of upstream files: + +| File in this folder | File in the upstream project | +| --- | --- | +| `SKILL.md` | `skills/simple-english/SKILL.md` | +| `references/rule-catalog.md` | `skills/simple-english/references/rule-catalog.md` | +| `references/strict-vocabulary.md` | `skills/simple-english/references/strict-vocabulary.md` | +| `references/use-cases.md` | `skills/simple-english/references/use-cases.md` | +| `references/word-swaps.md` | `skills/simple-english/references/word-swaps.md` | +| `references/system-prompt.md` | `prompts/system-prompt.md` | +| `scripts/ste_lint.py` | `evals/ste_lint.py` | +| `scripts/slop.tsv` | `evals/slop.tsv` | +| `LICENSE` | `LICENSE` | + +One copy has a change. In `scripts/ste_lint.py`, the `strip_code` function keeps the newlines that it removes. In the upstream version, each line number after a code block or a table is too small. The hook compares these line numbers with `git diff`, so the numbers must be correct. The change does not change the violation counts. + +## Files for this repository + +`scripts/hook.py` is the hook script for Cursor, Claude Code, and Codex. It replaces the two upstream hook scripts, `src/hooks/simple-english-activate.js` and `src/hooks/lint_hook.py`. The upstream scripts read the layout of a plugin, and they support Claude Code and Codex only. `scripts/test_hook.py` holds the tests for `scripts/hook.py`. + +## Update the copy + +1. Copy the upstream files in the table above into this folder. +2. If the upstream `strip_code` function does not keep newlines, apply the change to `scripts/ste_lint.py` again. +3. Change the commit at the top of this file. +4. Run `python3 .agents/skills/simple-english/scripts/test_hook.py`. +5. Run `python3 .agents/skills/simple-english/scripts/ste_lint.py --self-test`. diff --git a/.agents/skills/simple-english/scripts/hook.py b/.agents/skills/simple-english/scripts/hook.py new file mode 100644 index 00000000..d1c92003 --- /dev/null +++ b/.agents/skills/simple-english/scripts/hook.py @@ -0,0 +1,255 @@ +#!/usr/bin/env python3 +"""Simple English hooks for Cursor, Claude Code, and Codex. + +The hooks are advisory. They never block an action, and a failure in this +script never stops a session. + +Usage: hook.py EVENT --harness HARNESS + + session-start Send the Simple English rules to the agent. + post-edit Lint a Markdown file that the agent wrote. + stop Check the last reply for formatting (Claude Code only). + +HARNESS is cursor, claude, or codex. It selects the input and output format. + +The lint reports only the lines that differ from the last commit, so that an +agent does not rewrite text that its task did not touch. A new file gets a full +lint. + +Set SIMPLE_ENGLISH_HOOKS=off to turn off all the hooks. The tests set +SIMPLE_ENGLISH_REPO_ROOT to use a different repository. +""" +import argparse +import fnmatch +import json +import os +import pathlib +import re +import subprocess +import sys + +HERE = pathlib.Path(__file__).resolve().parent +SKILL_DIR = HERE.parent +REPO_ROOT = pathlib.Path(os.environ.get("SIMPLE_ENGLISH_REPO_ROOT") or SKILL_DIR.parent.parent.parent) +RULES_FILE = SKILL_DIR / "references" / "system-prompt.md" +sys.path.insert(0, str(HERE)) +# The hook runs in every checkout. Do not leave __pycache__ folders in the working tree. +sys.dont_write_bytecode = True + +# Claude Code writes hook output over 10,000 characters to a file and sends only a preview. +MAX_CONTEXT_CHARS = 9500 +MAX_LINT_HITS = 12 + +# Paths relative to the repository root. The file check does not lint these files. +SKIP_PATTERNS = ( + ".agents/skills/simple-english/*", + "CHANGELOG.md", + "*/node_modules/*", + "vendor/*", +) + +HEADER = """SIMPLE ENGLISH RULES FOR THIS REPOSITORY + +Write all prose in Simple English: replies, Markdown files, pull request descriptions, commit messages, and code comments. The rules follow. The full skill, with the rule catalog and the check mode, is at .agents/skills/simple-english/SKILL.md. Read it before you write or rewrite a document. + +""" + +FALLBACK_RULES = ( + "Use short sentences, active voice, and simple tenses. Put a condition before its command. " + "Use one word for one meaning. Do not change code, identifiers, commands, or quoted errors." +) + +OPENERS = re.compile(r"^\s*(certainly|great question|you're absolutely right|sure[,!]|absolutely[,!])", re.I) +CLOSERS = re.compile(r"(i hope this helps|let me know if|feel free to)", re.I) + + +def rule_block(text): + """The rules between the first two "---" lines of the prompt file.""" + parts = re.split(r"^---[ \t]*$", text, flags=re.M) + return parts[1].strip() if len(parts) >= 3 else text.strip() + + +def session_context(): + try: + rules = rule_block(RULES_FILE.read_text(encoding="utf-8")) + except OSError: + rules = FALLBACK_RULES + context = HEADER + rules + if len(context) > MAX_CONTEXT_CHARS: + context = HEADER + FALLBACK_RULES + return context + + +def load_linter(): + try: + import ste_lint # noqa: WPS433 + + return ste_lint + except Exception: # noqa: BLE001 + return None + + +def called_by_cursor(event): + """Cursor also runs the hooks in .claude/settings.json. Cursor runs its own copy from .cursor/hooks.json.""" + return bool(event.get("cursor_version") or os.environ.get("CURSOR_VERSION")) + + +def edited_path(event): + tool_input = event.get("tool_input") or {} + if isinstance(tool_input, str): + try: + tool_input = json.loads(tool_input) + except ValueError: + return None + if not isinstance(tool_input, dict): + return None + for key in ("file_path", "path", "target_file", "filePath"): + value = tool_input.get(key) + if isinstance(value, str) and value: + return value + return event.get("file_path") or None + + +def lint_target(event, raw_path): + """The absolute path of the Markdown file to lint, or None to skip the file.""" + if not raw_path or not raw_path.endswith(".md"): + return None + base = event.get("cwd") or os.environ.get("CURSOR_PROJECT_DIR") or os.environ.get("CLAUDE_PROJECT_DIR") or os.getcwd() + target = pathlib.Path(base, pathlib.Path(raw_path).expanduser()).resolve() + try: + relative = target.relative_to(REPO_ROOT.resolve()).as_posix() + except ValueError: + return None + if any(fnmatch.fnmatch(relative, pattern) for pattern in SKIP_PATTERNS): + return None + return target + + +HUNK = re.compile(r"^@@ -\S+ \+(\d+)(?:,(\d+))? @@", re.M) + + +def git(*args): + return subprocess.run(["git", "-C", str(REPO_ROOT), *args], capture_output=True, text=True, timeout=5) + + +def changed_lines(target): + """The line numbers that differ from the last commit. None means all lines.""" + try: + diff = git("diff", "--no-color", "--unified=0", "HEAD", "--", str(target)) + if diff.returncode != 0: + return None + if not diff.stdout: + tracked = git("ls-files", "--error-unmatch", "--", str(target)) + return set() if tracked.returncode == 0 else None + except (OSError, subprocess.SubprocessError): + return None + lines = set() + for match in HUNK.finditer(diff.stdout): + start, count = int(match.group(1)), int(match.group(2) or 1) + lines.update(range(start, start + count)) + return lines + + +def lint_report(target): + """A short report of the violations in the changed lines, or None when they have none.""" + lint = load_linter() + if lint is None: + return None + try: + text = target.read_text(encoding="utf-8") + except OSError: + return None + hits = lint.lint_detail(text, "descriptive") + changed = changed_lines(target) + if changed is not None: + hits = [hit for hit in hits if hit["line"] in changed] + if not hits: + return None + scope = "in the file" if changed is None else "in the lines that you changed" + lines = [f"simple-english: {target.name} has {len(hits)} Simple English violations {scope}."] + for hit in hits[:MAX_LINT_HITS]: + lines.append(f" line {hit['line']}, {hit['category']}: {hit['text']}") + if len(hits) > MAX_LINT_HITS: + lines.append(f" {len(hits) - MAX_LINT_HITS} more violations are not shown.") + lines.append("Correct these violations in the lines that you wrote. Do not change code, quoted text, or other lines.") + return "\n".join(lines) + + +def reply_problems(reply): + problems = [] + lint = load_linter() + if lint is not None: + counts = lint.reader_check(reply)["counts"] + for key, label in (("em_dash", "em dash"), ("bold_spans", "bold span"), ("headers", "header"), ("bullets", "list item")): + if counts[key]: + problems.append(f"{counts[key]} {label}(s)") + prose = re.sub(r"```.*?```", " ", reply, flags=re.S) + prose = re.sub(r"`[^`]*`", " ", prose) + slop = lint.lint(prose, "descriptive")["violations"].get("slop_word", 0) + if slop: + problems.append(f"{slop} slop word(s)") + if OPENERS.search(reply): + problems.append("a filler opener") + if CLOSERS.search(reply): + problems.append("a filler closer") + return problems + + +def session_start(harness, event): + context = session_context() + if harness == "cursor": + print(json.dumps({"additional_context": context})) + else: + sys.stdout.write(context) + return 0 + + +def post_edit(harness, event): + target = lint_target(event, edited_path(event)) + report = lint_report(target) if target else None + if not report: + return 0 + if harness == "cursor": + print(json.dumps({"additional_context": report})) + return 0 + # Claude Code shows stderr to the model when a PostToolUse hook exits with 2. The tool already ran. + sys.stderr.write(report + "\n") + return 2 + + +def stop(harness, event): + problems = reply_problems(event.get("last_assistant_message") or "") + if problems: + message = "simple-english reply check: " + ", ".join(problems) + ". Answer in prose." + print(json.dumps({"systemMessage": message})) + return 0 + + +EVENTS = {"session-start": session_start, "post-edit": post_edit, "stop": stop} + + +def main(argv=None): + parser = argparse.ArgumentParser(description="Simple English hooks.") + parser.add_argument("event", choices=sorted(EVENTS)) + parser.add_argument("--harness", choices=("cursor", "claude", "codex"), required=True) + args = parser.parse_args(argv) + + if os.environ.get("SIMPLE_ENGLISH_HOOKS", "").lower() == "off": + return 0 + try: + raw = sys.stdin.read() if not sys.stdin.isatty() else "" + event = json.loads(raw) if raw.strip() else {} + except ValueError: + event = {} + if not isinstance(event, dict): + event = {} + if args.harness == "claude" and called_by_cursor(event): + return 0 + try: + return EVENTS[args.event](args.harness, event) + except Exception: # noqa: BLE001 An advisory hook must never block or loop a session. + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/.agents/skills/simple-english/scripts/test_hook.py b/.agents/skills/simple-english/scripts/test_hook.py new file mode 100644 index 00000000..cf5ad7f3 --- /dev/null +++ b/.agents/skills/simple-english/scripts/test_hook.py @@ -0,0 +1,179 @@ +#!/usr/bin/env python3 +"""Tests for hook.py. Run: python3 .agents/skills/simple-english/scripts/test_hook.py""" +import json +import os +import pathlib +import subprocess +import sys +import tempfile +import unittest + +HOOK = pathlib.Path(__file__).resolve().parent / "hook.py" +REPO_ROOT = HOOK.parent.parent.parent.parent.parent + +BAD_MARKDOWN = "You should run the command; it's important.\n" +GOOD_MARKDOWN = "Run the command. The command deletes old rows.\n" + + +def run(event_name, harness, payload=None, env=None): + full_env = {k: v for k, v in os.environ.items() if not k.startswith(("CURSOR_", "SIMPLE_ENGLISH_"))} + full_env.update(env or {}) + stdin = "" if payload is None else (payload if isinstance(payload, str) else json.dumps(payload)) + return subprocess.run( + [sys.executable, str(HOOK), event_name, "--harness", harness], + input=stdin, + capture_output=True, + text=True, + env=full_env, + cwd=REPO_ROOT, + timeout=30, + ) + + +class MarkdownFile: + """A temporary Markdown file inside the repository, because the hook skips files outside it.""" + + def __init__(self, text): + self.text = text + + def __enter__(self): + self.dir = tempfile.TemporaryDirectory(dir=REPO_ROOT, prefix="simple-english-test-") + self.path = pathlib.Path(self.dir.name, "doc.md") + self.path.write_text(self.text, encoding="utf-8") + return self.path + + def __exit__(self, *exc): + self.dir.cleanup() + + +class SessionStartTest(unittest.TestCase): + def test_cursor_gets_additional_context_json(self): + result = run("session-start", "cursor", {"hook_event_name": "sessionStart"}) + self.assertEqual(result.returncode, 0) + context = json.loads(result.stdout)["additional_context"] + self.assertIn("SIMPLE ENGLISH RULES FOR THIS REPOSITORY", context) + self.assertIn("THE REPLY", context) + self.assertLess(len(context), 9500) + + def test_claude_and_codex_get_plain_text(self): + for harness in ("claude", "codex"): + result = run("session-start", harness, {"hook_event_name": "SessionStart"}) + self.assertEqual(result.returncode, 0) + self.assertTrue(result.stdout.startswith("SIMPLE ENGLISH RULES FOR THIS REPOSITORY")) + self.assertIn("THE DOCUMENT", result.stdout) + + def test_claude_hook_does_nothing_when_cursor_runs_it(self): + from_stdin = run("session-start", "claude", {"hook_event_name": "sessionStart", "cursor_version": "2.0.0"}) + from_env = run("session-start", "claude", {}, env={"CURSOR_VERSION": "2.0.0"}) + for result in (from_stdin, from_env): + self.assertEqual(result.returncode, 0) + self.assertEqual(result.stdout, "") + + def test_off_switch(self): + result = run("session-start", "cursor", {}, env={"SIMPLE_ENGLISH_HOOKS": "off"}) + self.assertEqual(result.returncode, 0) + self.assertEqual(result.stdout, "") + + +class PostEditTest(unittest.TestCase): + def test_cursor_reports_violations_as_additional_context(self): + with MarkdownFile(BAD_MARKDOWN) as path: + result = run("post-edit", "cursor", {"tool_name": "Write", "tool_input": {"file_path": str(path)}}) + self.assertEqual(result.returncode, 0) + report = json.loads(result.stdout)["additional_context"] + self.assertIn("doc.md has", report) + self.assertIn("banned_modal", report) + + def test_cursor_accepts_a_relative_path_and_a_string_tool_input(self): + with MarkdownFile(BAD_MARKDOWN) as path: + relative = path.relative_to(REPO_ROOT).as_posix() + result = run("post-edit", "cursor", {"tool_input": json.dumps({"path": relative}), "cwd": str(REPO_ROOT)}) + self.assertIn("additional_context", json.loads(result.stdout)) + + def test_claude_reports_violations_on_stderr_with_exit_2(self): + with MarkdownFile(BAD_MARKDOWN) as path: + result = run("post-edit", "claude", {"hook_event_name": "PostToolUse", "tool_input": {"file_path": str(path)}}) + self.assertEqual(result.returncode, 2) + self.assertIn("contraction", result.stderr) + self.assertEqual(result.stdout, "") + + def test_clean_file_gives_no_output(self): + with MarkdownFile(GOOD_MARKDOWN) as path: + result = run("post-edit", "cursor", {"tool_input": {"file_path": str(path)}}) + self.assertEqual((result.returncode, result.stdout), (0, "")) + + def test_skips_files_that_are_not_linted(self): + skipped = [ + REPO_ROOT / "main.go", + REPO_ROOT / "CHANGELOG.md", + REPO_ROOT / ".agents/skills/simple-english/SKILL.md", + pathlib.Path(tempfile.gettempdir(), "outside-the-repository.md"), + ] + for path in skipped: + result = run("post-edit", "cursor", {"tool_input": {"file_path": str(path)}}) + self.assertEqual((result.returncode, result.stdout), (0, ""), path) + + +class ChangedLinesTest(unittest.TestCase): + """For a committed file, the lint reports only the lines that differ from the last commit.""" + + def setUp(self): + self.dir = tempfile.TemporaryDirectory(prefix="simple-english-repo-") + self.root = pathlib.Path(self.dir.name) + self.doc = self.root / "doc.md" + self.doc.write_text(BAD_MARKDOWN, encoding="utf-8") + git = ["git", "-C", str(self.root), "-c", "user.name=test", "-c", "user.email=test@example.com"] + subprocess.run([*git, "init", "-q"], check=True) + subprocess.run([*git, "add", "doc.md"], check=True) + subprocess.run([*git, "commit", "-q", "-m", "Add doc"], check=True) + self.env = {"SIMPLE_ENGLISH_REPO_ROOT": str(self.root)} + + def tearDown(self): + self.dir.cleanup() + + def edit(self, text): + self.doc.write_text(BAD_MARKDOWN + text, encoding="utf-8") + return run("post-edit", "cursor", {"tool_input": {"file_path": str(self.doc)}}, env=self.env) + + def test_ignores_violations_in_lines_that_did_not_change(self): + result = self.edit(GOOD_MARKDOWN) + self.assertEqual((result.returncode, result.stdout), (0, "")) + + def test_reports_violations_in_a_changed_line(self): + result = self.edit("You'll see that it has been removed.\n") + report = json.loads(result.stdout)["additional_context"] + self.assertIn("in the lines that you changed", report) + self.assertIn("line 2,", report) + self.assertNotIn("line 1,", report) + + def test_line_numbers_are_correct_after_code_blocks_and_tables(self): + blocks = "\n```bash\nmake build\nmake test\n```\n\n| Name | Value |\n| --- | --- |\n| a | b |\n\n" + result = self.edit(blocks + "You should not do this.\n") + report = json.loads(result.stdout)["additional_context"] + bad_line = (BAD_MARKDOWN + blocks).count("\n") + 1 + self.assertIn(f"line {bad_line}, banned_modal: should", report) + + +class StopTest(unittest.TestCase): + def test_reports_formatting_in_a_reply(self): + result = run("stop", "claude", {"hook_event_name": "Stop", "last_assistant_message": "**Yes** — it works.\n\n- one\n- two\n"}) + self.assertEqual(result.returncode, 0) + message = json.loads(result.stdout)["systemMessage"] + self.assertIn("bold span", message) + self.assertIn("em dash", message) + + def test_prose_reply_gives_no_output(self): + result = run("stop", "claude", {"last_assistant_message": "The build passed. You can merge the change."}) + self.assertEqual((result.returncode, result.stdout), (0, "")) + + +class BadInputTest(unittest.TestCase): + def test_malformed_or_missing_input_never_fails(self): + for event_name in ("session-start", "post-edit", "stop"): + for payload in ("not json", "[]", ""): + result = run(event_name, "claude", payload) + self.assertEqual(result.returncode, 0, (event_name, payload)) + + +if __name__ == "__main__": + unittest.main() diff --git a/.claude/settings.json b/.claude/settings.json new file mode 100644 index 00000000..8f80646e --- /dev/null +++ b/.claude/settings.json @@ -0,0 +1,38 @@ +{ + "hooks": { + "SessionStart": [ + { + "hooks": [ + { + "type": "command", + "command": "python3 \"$CLAUDE_PROJECT_DIR/.agents/skills/simple-english/scripts/hook.py\" session-start --harness claude", + "timeout": 5 + } + ] + } + ], + "PostToolUse": [ + { + "matcher": "Write|Edit", + "hooks": [ + { + "type": "command", + "command": "python3 \"$CLAUDE_PROJECT_DIR/.agents/skills/simple-english/scripts/hook.py\" post-edit --harness claude", + "timeout": 10 + } + ] + } + ], + "Stop": [ + { + "hooks": [ + { + "type": "command", + "command": "python3 \"$CLAUDE_PROJECT_DIR/.agents/skills/simple-english/scripts/hook.py\" stop --harness claude", + "timeout": 10 + } + ] + } + ] + } +} diff --git a/.codex/hooks.json b/.codex/hooks.json new file mode 100644 index 00000000..c29cbaeb --- /dev/null +++ b/.codex/hooks.json @@ -0,0 +1,19 @@ +{ + "description": "Load the Simple English rules at the start of each Codex session.", + "hooks": { + "SessionStart": [ + { + "matcher": "^(startup|resume|clear|compact)$", + "hooks": [ + { + "type": "command", + "command": "python3 \"$(git rev-parse --show-toplevel)/.agents/skills/simple-english/scripts/hook.py\" session-start --harness codex", + "timeout": 5, + "statusMessage": "Loading Simple English...", + "additionalContextLimit": 0 + } + ] + } + ] + } +} diff --git a/.cursor/hooks.json b/.cursor/hooks.json new file mode 100644 index 00000000..bff03c95 --- /dev/null +++ b/.cursor/hooks.json @@ -0,0 +1,18 @@ +{ + "version": 1, + "hooks": { + "sessionStart": [ + { + "command": "python3 .agents/skills/simple-english/scripts/hook.py session-start --harness cursor", + "timeout": 5 + } + ], + "postToolUse": [ + { + "command": "python3 .agents/skills/simple-english/scripts/hook.py post-edit --harness cursor", + "matcher": "Write", + "timeout": 10 + } + ] + } +} From 1ac414937961f483976f83322c64d1e368df9528 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Tue, 29 Sep 2026 21:33:08 +0000 Subject: [PATCH 4/5] docs: tell agents to write in Simple English Add a Writing Style section to AGENTS.md. CLAUDE.md is a link to this file. The section gives the rules that agents break most often, the text that the rules cover, and the hooks for each tool. Co-authored-by: Ramon Niebla --- AGENTS.md | 50 ++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 50 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index aeac8d98..8e6787a2 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -6,6 +6,56 @@ This file provides guidance to AI coding agents when working with code in this r LaunchDarkly CLI (`ldcli`) — a Go CLI for managing LaunchDarkly feature flags. Built with Cobra/Viper, distributed via Homebrew, Docker, NPM, and GitHub Releases. +## Writing Style: Simple English + +Write all prose in Simple English. Simple English is plain English that follows the rules of ASD-STE100 Simplified Technical English. The full rules are in [`.agents/skills/simple-english/SKILL.md`](.agents/skills/simple-english/SKILL.md). Read that file before you write or rewrite a document. + +The rules apply to this text: + +- Replies to the user +- Markdown files, such as `README.md` and `CONTRIBUTING.md` +- Pull request titles and descriptions +- Commit messages +- New code comments, CLI help text, and error messages + +Do not change these items: + +- Code, identifiers, commands, flags, file paths, and quoted errors +- Generated files, such as `CHANGELOG.md` and `cmd/resources/resource_cmds.go` +- Text that your task does not touch + +Agents break these rules most often: + +1. Write short sentences. Use 20 words at most for an instruction and 25 words at most for a description. +2. Use active voice and simple tenses. Write `The command deleted the row`, not `The row has been deleted`. +3. Use `can`, `will`, or `must`. Do not use `should`, `would`, `may`, `might`, or `could`. +4. Put a condition before its command: "If the build fails, read the log." +5. Use one word for one meaning in a document. For example, use `configuration` every time, not `config` in one place and `settings` in another. +6. Define a technical term the first time that you use it. +7. Do not use contractions, semicolons, or em dashes. +8. State facts. Do not add words such as `robust`, `seamless`, or `crucial`. +9. In a reply, put the answer in the first sentence. Write prose, with no headers, bold text, lists, or tables. + +### Agent Hooks + +Hooks load these rules at the start of a session. They also lint the Markdown files that an agent writes. The hooks are advisory, so they never block an action. Each hook runs `.agents/skills/simple-english/scripts/hook.py`, which needs `python3`. + +| Tool | Configuration | What the hooks do | +| --- | --- | --- | +| Cursor | `.cursor/hooks.json` | Load the rules at session start. Lint each Markdown file after a write. | +| Claude Code | `.claude/settings.json` | Load the rules at session start. Lint each Markdown file after a write. Report bold text, headers, lists, and em dashes in each reply. | +| Codex | `.codex/hooks.json` | Load the rules at session start. | + +These limits apply: + +- Cursor Cloud Agents do not run session start hooks. They get the rules from this file, and the Markdown lint still runs. +- Cursor also runs the hooks in `.claude/settings.json`. The script finds this case and runs only the Cursor hooks. +- Codex runs project hooks only in a trusted project. Codex also asks you to approve each hook. Open `/hooks` to approve it. +- The lint reports only the lines that differ from the last commit. A new file gets a full lint. +- The lint skips `CHANGELOG.md`, the skill folder, and files outside the repository. + +To turn off the hooks, set `SIMPLE_ENGLISH_HOOKS=off`. To test the hooks, run `python3 .agents/skills/simple-english/scripts/test_hook.py`. The file `.agents/skills/simple-english/UPSTREAM.md` gives the source of the skill and the steps to update it. + ## Common Commands ```bash From e8c770adad3a823f66f350c964df480eb0ebfec5 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Tue, 29 Sep 2026 21:37:36 +0000 Subject: [PATCH 5/5] fix: remove the extra blank line at the end of rule-catalog.md The upstream file ends with a blank line. The end-of-file-fixer pre-commit hook requires one newline at the end of each file, so the build failed. UPSTREAM.md now records this change. Co-authored-by: Ramon Niebla --- .agents/skills/simple-english/UPSTREAM.md | 7 +++++-- .agents/skills/simple-english/references/rule-catalog.md | 1 - 2 files changed, 5 insertions(+), 3 deletions(-) diff --git a/.agents/skills/simple-english/UPSTREAM.md b/.agents/skills/simple-english/UPSTREAM.md index 9bac87ce..bf9e30b3 100644 --- a/.agents/skills/simple-english/UPSTREAM.md +++ b/.agents/skills/simple-english/UPSTREAM.md @@ -18,7 +18,10 @@ These files are copies of upstream files: | `scripts/slop.tsv` | `evals/slop.tsv` | | `LICENSE` | `LICENSE` | -One copy has a change. In `scripts/ste_lint.py`, the `strip_code` function keeps the newlines that it removes. In the upstream version, each line number after a code block or a table is too small. The hook compares these line numbers with `git diff`, so the numbers must be correct. The change does not change the violation counts. +Two copies have a change: + +- In `scripts/ste_lint.py`, the `strip_code` function keeps the newlines that it removes. In the upstream version, each line number after a code block or a table is too small. The hook compares these line numbers with `git diff`, so the numbers must be correct. The change does not change the violation counts. +- In `references/rule-catalog.md`, the blank line at the end of the file is removed. The `end-of-file-fixer` pre-commit hook of this repository requires one newline at the end of each file. ## Files for this repository @@ -27,7 +30,7 @@ One copy has a change. In `scripts/ste_lint.py`, the `strip_code` function keeps ## Update the copy 1. Copy the upstream files in the table above into this folder. -2. If the upstream `strip_code` function does not keep newlines, apply the change to `scripts/ste_lint.py` again. +2. Apply the two changes above again, if the upstream files still need them. 3. Change the commit at the top of this file. 4. Run `python3 .agents/skills/simple-english/scripts/test_hook.py`. 5. Run `python3 .agents/skills/simple-english/scripts/ste_lint.py --self-test`. diff --git a/.agents/skills/simple-english/references/rule-catalog.md b/.agents/skills/simple-english/references/rule-catalog.md index d2356fb1..9f7fb432 100644 --- a/.agents/skills/simple-english/references/rule-catalog.md +++ b/.agents/skills/simple-english/references/rule-catalog.md @@ -162,4 +162,3 @@ One word, one meaning, one part of speech, for the whole document (Rules 1.11, 9 - The settings file is `configuration`, never config, settings, or options in the same document. - The verify concept is `make sure that`, never check, verify, confirm, validate, or ensure as verbs. Strict mode routes the rest with `references/strict-vocabulary.md`. - Common swaps: however → but, therefore → as a result, since (= because) → because, perform → do, avoid → prevent, repeat → do again, acceptable → permitted, now → delete it. -