Skip to content

Landing page: score a prompt in the browser; README rewrite with an architecture diagram - #25

Merged
Bogzx merged 14 commits into
mainfrom
improve/2026-10-01-try-it-docs
Oct 1, 2026
Merged

Bogzx merged 14 commits into
mainfrom
improve/2026-10-01-try-it-docs

Conversation

@Bogzx

@Bogzx Bogzx commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Last in the stack #22 → #23 → #24 → this. Merge those first; this PR's own commits are the last two. It also merges cleanly with #21 (landing-page honesty), checked with a test merge: the two touch different files.

Why

  • Nothing to try. The landing page is the project's public "demo" link, and there was nothing on it to try.
  • The README was 452 lines and out of date. It still described the hackathon deployment: Postgres on Neon with sslmode=require, the API on Railway, team tokens "derived from the git remote" (replaced on 2026-09-30), and an undefined "HCL bundle". It had no picture of how the pieces connect.

What changed

  1. Try it (8b7dba0). A section under the hero scores whatever you type, live, with the rule-based scorer from Keyless demo mode, a measured rule-based scorer, seeded demo data, demo-team switch, compose smoke test #23. It shows five bars, the missing-dimension hints, and the overall score with what LearnLoop would do at that score. It offers three example prompts.

    • It says it is the rule-based scorer, not the model, and links to how both are measured. Nothing typed leaves the page.
    • The page has no build step, so it carries a byte-identical copy of heuristic-score.mjs in assets/heuristic-score.js. A test in packages/scoring fails if the copy drifts.
  2. README (c645348). In order:

    • what it is and who built it (the four GitHub contributors);
    • three ways to try it: the browser, TRAILHEAD_LLM=offline docker compose --profile demo up, or with a Gemini key;
    • a Mermaid architecture diagram and the five-step loop;
    • the rubric, with what has and hasn't been measured;
    • a surface table, development, deployment, specs, licence.

    The route table moved to apps/api/README.md and the full env table to SELFHOSTING.md → All variables (DATABASE_URL/PORT rows corrected). New CHANGELOG.md for the v0.1.0 tag.

Verification

  • Headless Chromium, landing page served locally. The examples score 0 / 3 / 10, a typed prompt scores 8, and the 1280 px and 390 px layouts render without horizontal scroll. With Landing page: replace the fake waitlist with real links, fix claims the code doesn't back #21 test-merged in: no waitlist, the hero links render, and the scorer works.
  • Copy-sync test. Appending a line to the copy turns it red; restoring turns it green.
  • Mermaid block. Renders with mermaid 11 (10 nodes, no errors).
  • Claims. Every model, route, env and behaviour claim in the README was read against the code. Claims that couldn't be verified were dropped (e.g. "built in 48 hours"). The releases link says packages arrive "from v0.1.0", because no release exists yet.
  • Checks. npm run typecheck, npm run lint and npm test (scoring 44 tests, +1 sync test) all pass.

After merging

🤖 Generated with Claude Code

Bogdan Truta and others added 14 commits October 1, 2026 12:26
app.ts held every route, its helpers and the middleware in one 87 KB file.
It is now the app shell (CORS, body caps, tracing, auth, LLM rate limits,
GET /, route mounting, error handler: 243 lines), with:

- http.ts: what the middleware and the routes share (env type, tenant
  policy, rate limiting, the prompt cap, skill-observation writes)
- routes/{score,coach,assist,wiki,prompts,metrics,teams,onboard}.ts: one
  Hono sub-app per area, mounted after the middleware so every route still
  gets auth, body caps, tracing and rate limits

Code moved verbatim (apart from `app.<verb>` → `<router>.<verb>` and
`export` on the shared helpers), so behaviour is unchanged: the 38
Postgres integration tests, which drive every route and cross-check the
GET / catalog, pass unchanged. apps/api/README.md documents the layout.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
wiki-bootstrap-job.ts built its own GoogleGenAI client, so the rich
bootstrap (the biggest fan-out of LLM calls in the API) bypassed the retry
policy and never appeared in Langfuse traces. It now calls
gemini.generateText, which goes through the same tracedGenerate/withRetry
path as the request handlers: 429/5xx are retried with backoff, timeouts
are still never retried (90 s per call, as before). generateText does not
apply the 4 000-char cap meant for short structured answers, because folder
narratives are several thousand words; a new test pins both properties.

Removed as dead code: extractLearning (gemini.ts), EXTRACT_MODEL
(gemma-4-31b-it) and the extract-prompt module in packages/scoring with its
tests. They served the Stop hook that was removed in PR #5; nothing has
imported them since. The README's "Gemma for async learning extraction"
claim goes with them; every call uses gemini-3-flash-preview.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ecutable

Next 16.3 generates AGENTS.md and CLAUDE.md in apps/dashboard on every
`next dev` unless `agentRules: false` is set; they showed up as untracked
files for anyone running the dashboard. apps/mcp-server/bin/cli.mjs has a
shebang and is the package's `bin`, but was committed as 100644 (npm fixed
the mode on install, leaving a dirty tree).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…utes/

The README said gemma-4-31b-it handled "async learning extraction" and
"diff narration, rich bootstrap". Learning extraction was never wired up
(removed in the previous commits) and /diff and the bootstrap use
gemini-3-flash-preview. Also points the route table at the new layout.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…l that runs in CI

The scorer at the centre of the product (one Gemini call) had never been
measured: the eval harness existed but needs a key, so nothing ran. This
adds the part that can run anywhere.

packages/scoring/src/heuristic-score.mjs scores a prompt on the five rubric
dimensions with transparent rules over surface features (file paths,
identifiers, numbers with units, constraint and output phrasing, meta
instructions aimed at the scorer). It is deterministic, has no imports and
no Node APIs, so the API's offline mode and the landing page can both use
it. It is a floor for the model, not a substitute.

The eval gains:
- an ordering metric: for every pair of prompts whose expected bands don't
  overlap, does the scorer rank the better one higher (ties count as wrong);
- `--scorer heuristic` (free, no key, no --yes);
- eval/holdout.json: 16 prompts with bands, written before the rule set was
  frozen and never used to tune it (golden.json was visible while writing);
- baseline.test.ts: runs the rule-based scorer over both sets in `npm test`
  and fails if it drops below what it measured when frozen.

Measured (heuristic-v1):
  golden  (seen):   29/30 overall in band, 47/49 dims, 152/152 pairs ordered
  holdout (unseen): 14/16 overall in band, 18/22 dims,   38/38 pairs ordered
The known misses are listed in eval/README.md. The Gemini scorer still has
not been run on either set; that needs a key.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nothing ran without a Gemini key: the API refused to boot, so a stranger
could not see the coach work, and nobody could try the stack before getting
a key.

TRAILHEAD_LLM=offline runs the same API with no model:
- llm.ts: every LLM call the routes make goes through it and picks its
  provider per call (llm-mode.ts): gemini.ts by default, offline-llm.ts in
  offline mode. gemini.ts now creates its client on first use, so importing
  it needs no key.
- offline-llm.ts: the rule-based scorer from packages/scoring for /score and
  /coach; keyword rules for topics; '' for the LLM-written rewrites, tips,
  acknowledgements and summaries (the /coach renderers already fall back to
  their static templates); a template /diff narrative naming the biggest
  gap; an /improve that asks the static question per weak dimension and
  appends the answers. POST /onboard/repo/full answers 503 llm_unavailable.
- It says so everywhere: GET / reports "llm", /score and /coach return
  "scorer": "heuristic", coaching text ends with a note, the score card in
  the extensions shows "Rule-based score", and the API warns at startup.
- It never switches on by itself: with no key and no TRAILHEAD_LLM the API
  still refuses to start (now naming both ways to fix it), so a deploy that
  lost its key fails loudly instead of quietly scoring with rules.

docker-compose.yml no longer requires the key up front, so
`TRAILHEAD_LLM=offline docker compose up` works with no .env at all.
SELFHOSTING.md documents what offline mode does and does not do, including
that the confirming re-score for library promotion is the same rules.

Tests: llm-mode and offline-llm unit tests (with fetch stubbed to fail), and
an integration test against real Postgres that scores, coaches, promotes,
diffs and improves in offline mode and asserts no request-path Gemini call
was made (verified: pointing /score back at gemini.ts makes it fail).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…DEMO_TEAM)

The seeded demo team's secret (trailhead_demo_acme_2026) is in this repo, and
schema.sql always creates the team. On a server reachable from a network,
anyone could write to its wiki and spend the operator's Gemini quota through
it; the only bound was the per-team and per-IP rate limits, and there was no
way to turn it off.

TRAILHEAD_DEMO_TEAM=on|off controls it. The default stays on for a local
setup, but setting TRAILHEAD_ADMIN_TOKEN (the sign of a networked deploy)
turns it off unless TRAILHEAD_DEMO_TEAM=on says otherwise. Off, the demo
secret gets 401 with reason demo_team_disabled and a message saying what to
do; the team row stays, so turning it back on just works. GET / reports
`demo_team`.

Integration test: default on; off once an admin token is set (other teams
unaffected); an explicit value wins in both directions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…demo up`

The seed wrote a wiki and four library prompts but no prompt activity
("Full §11 seed ... deferred"), and compose never ran it, so a fresh stack
showed an empty skill-arc chart and zeros on the team page.

packages/db/seed.mjs now also writes, for three made-up users, 48 scored
prompts each spread over the last six days with scores that drift upwards
(720 skill_observations, deterministic: fixed-seed PRNG), plus 18 captures
with outcomes. It stays idempotent: the synthetic rows are replaced on every
run, which also moves the window up to "now" (the dashboard shows the last
24 h and the team page the last 7 days). The header says plainly that all of
it is synthetic, and the dashboard's demo-team card no longer calls the
fictional Acme Fintech "the team's actual repo conventions".

docker-compose.yml gains a one-shot `seed` service under the `demo`
profile, started after the API is healthy:
  TRAILHEAD_LLM=offline docker compose --profile demo up --build

Verified with that command on a fresh volume: seed exits 0, /team/metrics
reports 720 observations, 3 active users, avg 5.2; the dashboard's
/skill-arc renders 24 hour-buckets and /team shows non-zero cards (headless
Chromium). Running the seed twice leaves the same counts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI built the API image but never ran the stack, so "docker compose up"
could break (schema mount, healthcheck, env wiring) without a red build.
The docker job now brings the stack up the way a stranger would, with no
key (TRAILHEAD_LLM=offline), runs the demo seed, and checks: GET / is ok
and offline, a team registers, a wiki write lands, /score answers with the
rule-based scorer, and the seeded demo team has its activity. Logs are
printed on failure and the stack is always torn down (down -v).

The same commands were run locally against a compose stack before
committing (all four jq checks true).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and gitignores what it creates

`init` wrote absolute paths into the conventionally shared project configs:
TRAILHEAD_TEAM_FILE=/home/<you>/repo/.trailhead-team and the path to the
server inside your LearnLoop clone. A teammate who pulled the repo got an
MCP config pointing at someone else's home directory.

- TRAILHEAD_TEAM_FILE is now the relative `.trailhead-team`. The server looks
  for a relative team file in its working directory, then each parent up to
  the git root (findTeamFile), so it works when a host starts it in a
  subdirectory and never reads a sentinel from another repo above the root.
  With --user-scope this also stops one absolute path from binding every
  project to the same team.
- The server entry still has to point at the clone (the package is not on
  npm), so a config `init` creates is added to .gitignore and `init` says
  why; a config that already existed (it may hold the team's other servers)
  is left alone, with a warning.

Verified: in a fresh git repo, `cli.mjs init` against a local API writes
`TRAILHEAD_TEAM_FILE: ".trailhead-team"` and gitignores .mcp.json; the MCP
server started from src/api/ with that env answered a `coach` call over
stdio. New tests for both behaviours; the CLI smoke tests now expect the
relative path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…verriding them

`init` writes the directive into the user's CLAUDE.md and
copilot-instructions.md, and the tool descriptions say the same. Both told
the agent that the tools "REPLACE native Read/Grep/Glob", that coach is
"MANDATORY before answering ANY code task" (an extra LLM round-trip on every
message), to save conventions "silently, never ask", and to start a rich
bootstrap on its own whenever the wiki looked empty — which uploads the
repo's source files to the LLM without the user having asked.

The directive (154 → 50 lines) and the tool descriptions now say:
- coach once at the start of each new code task, not on follow-ups; it
  never blocks work (the proceed/round_token/skip_reveal protocol is kept);
- wiki_lookup gives the team's conventions for an area, then read the code
  as usual;
- wiki_save when the user states a team rule, and say in one line that it
  was saved;
- wiki_bootstrap only when the user asks; if the wiki looks empty, offer it,
  and say that rich mode sends source files to the team's API.

A test pins the directive's shape (no REPLACE/INSTEAD OF/MANDATORY, no
unasked bootstrap, ≤ 60 lines). Verified over stdio: tools/list and
resources/read serve the new text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ed release

Neither extension could be installed without cloning and building: the
browser extension had no icons and no package, the VS Code .vsix shipped its
sources, tests, .vscode/ and a sourcemap (no .vscodeignore), and nothing
produced either file in CI.

- Browser extension: icons/icon.svg (the LearnLoop rings on the brand ink)
  rendered to 16/32/48/128 PNGs and declared in the manifest; the build
  copies them into dist/. `npm run package` makes a production build and
  learnloop-browser-ext-<version>.zip with a dependency-free zip writer
  (node:zlib deflate + crc32, fixed timestamps so the same dist zips to the
  same bytes). bundle-load.test checks every file the manifest references
  is in dist/.
- VS Code extension: .vscodeignore (sources, tests, maps out: 22 files /
  42 KB → 8 files / 18 KB), a marketplace icon, LICENSE in the package,
  displayName "LearnLoop", and `npm run package` (vsce 3.9.2 via npx).
- Both at 0.1.0 (lockfile entries patched to match).
- ci.yml: a `package` job builds both and uploads them as the `extensions`
  artifact on every run.
- release.yml: on a v* tag, checks the tag matches both manifests, tests,
  packages, and attaches both files to the release for that tag (creating
  it as a draft if needed). Nothing is published to any store.

Verified: the zip passes `unzip -t`; unzipped and loaded with
`chrome --load-extension`, Chromium started the extension's service worker
and its popup rendered. `vsce ls` lists only package.json, README, LICENSE,
media/ and dist/extension.js. The release workflow's version check accepts
v0.1.0 and rejects a mismatched tag (run locally with jq); the workflow
itself runs only when a tag is pushed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The landing page is the project's public demo link, and there was nothing on
it to try. A new section under the hero scores whatever you type, live, with
the same rule-based scorer the API uses in offline mode: five bars, the
missing-dimension hints, and the overall score with what LearnLoop would do
at that score. Three example prompts (vague / halfway / specific). It says
plainly that this is the rule-based scorer, not the model, and links to how
both are measured. Nothing typed leaves the page.

The page has no build step, so it carries a byte-identical copy of
packages/scoring/src/heuristic-score.mjs as assets/heuristic-score.js,
loaded with a module script; a test in packages/scoring fails if the copy
drifts (checked: appending a line turns it red).

Verified in headless Chromium against `python3 -m http.server`: the examples
score 0, 3 and 10, a typed prompt scores 8; desktop (1280) and phone (390)
layouts render without horizontal scroll. The section touches only new files
plus app.jsx and index.html, so it merges independently of #21.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…re diagram

The README was 452 lines and still described the hackathon setup: Postgres
on Neon with sslmode=require, the API on Railway, team tokens "derived from
the git remote" (replaced on 2026-09-30), an undefined "HCL bundle", and no
picture of how the pieces connect.

It now leads with what LearnLoop is, who built it (the four contributors on
GitHub), and three ways to try it: the in-browser scorer, the keyless
`docker compose --profile demo` stack, and the stack with a Gemini key. Then
a Mermaid diagram of clients → API → Postgres/LLM, the five-step loop, the
rubric with what has and has not been measured (the rule-based baseline's
numbers; the Gemini scorer not yet run), a surface table linking each app,
development, deployment, specs and licence.

Moved rather than dropped: the route table to apps/api/README.md, the full
environment-variable table to SELFHOSTING.md → All variables (with the
DATABASE_URL/PORT rows corrected). CHANGELOG.md starts at 0.1.0, the tag the
release workflow will package.

Checked: the Mermaid block renders with mermaid 11 in headless Chromium;
every model, route and env claim was read against the code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
learnloop Building Building Preview Oct 1, 2026 1:22pm UTC
polihackwinners-dashboard Ready Ready Preview Oct 1, 2026 1:22pm UTC

This branch was successfully deployed

2 active deployments
Preview – polihackwinners-dashboard — c6453488 Deployed Oct 1, 2026 by vercel[bot]
Preview – learnloop — c6453488 Deployed Oct 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant