Keyless demo mode, a measured rule-based scorer, seeded demo data, demo-team switch, compose smoke test - #23
Merged
Merged
Conversation
app.ts held every route, its helpers and the middleware in one 87 KB file.
It is now the app shell (CORS, body caps, tracing, auth, LLM rate limits,
GET /, route mounting, error handler: 243 lines), with:
- http.ts: what the middleware and the routes share (env type, tenant
policy, rate limiting, the prompt cap, skill-observation writes)
- routes/{score,coach,assist,wiki,prompts,metrics,teams,onboard}.ts: one
Hono sub-app per area, mounted after the middleware so every route still
gets auth, body caps, tracing and rate limits
Code moved verbatim (apart from `app.<verb>` → `<router>.<verb>` and
`export` on the shared helpers), so behaviour is unchanged: the 38
Postgres integration tests, which drive every route and cross-check the
GET / catalog, pass unchanged. apps/api/README.md documents the layout.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
wiki-bootstrap-job.ts built its own GoogleGenAI client, so the rich bootstrap (the biggest fan-out of LLM calls in the API) bypassed the retry policy and never appeared in Langfuse traces. It now calls gemini.generateText, which goes through the same tracedGenerate/withRetry path as the request handlers: 429/5xx are retried with backoff, timeouts are still never retried (90 s per call, as before). generateText does not apply the 4 000-char cap meant for short structured answers, because folder narratives are several thousand words; a new test pins both properties. Removed as dead code: extractLearning (gemini.ts), EXTRACT_MODEL (gemma-4-31b-it) and the extract-prompt module in packages/scoring with its tests. They served the Stop hook that was removed in PR #5; nothing has imported them since. The README's "Gemma for async learning extraction" claim goes with them; every call uses gemini-3-flash-preview. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ecutable Next 16.3 generates AGENTS.md and CLAUDE.md in apps/dashboard on every `next dev` unless `agentRules: false` is set; they showed up as untracked files for anyone running the dashboard. apps/mcp-server/bin/cli.mjs has a shebang and is the package's `bin`, but was committed as 100644 (npm fixed the mode on install, leaving a dirty tree). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…utes/ The README said gemma-4-31b-it handled "async learning extraction" and "diff narration, rich bootstrap". Learning extraction was never wired up (removed in the previous commits) and /diff and the bootstrap use gemini-3-flash-preview. Also points the route table at the new layout. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…l that runs in CI The scorer at the centre of the product (one Gemini call) had never been measured: the eval harness existed but needs a key, so nothing ran. This adds the part that can run anywhere. packages/scoring/src/heuristic-score.mjs scores a prompt on the five rubric dimensions with transparent rules over surface features (file paths, identifiers, numbers with units, constraint and output phrasing, meta instructions aimed at the scorer). It is deterministic, has no imports and no Node APIs, so the API's offline mode and the landing page can both use it. It is a floor for the model, not a substitute. The eval gains: - an ordering metric: for every pair of prompts whose expected bands don't overlap, does the scorer rank the better one higher (ties count as wrong); - `--scorer heuristic` (free, no key, no --yes); - eval/holdout.json: 16 prompts with bands, written before the rule set was frozen and never used to tune it (golden.json was visible while writing); - baseline.test.ts: runs the rule-based scorer over both sets in `npm test` and fails if it drops below what it measured when frozen. Measured (heuristic-v1): golden (seen): 29/30 overall in band, 47/49 dims, 152/152 pairs ordered holdout (unseen): 14/16 overall in band, 18/22 dims, 38/38 pairs ordered The known misses are listed in eval/README.md. The Gemini scorer still has not been run on either set; that needs a key. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nothing ran without a Gemini key: the API refused to boot, so a stranger could not see the coach work, and nobody could try the stack before getting a key. TRAILHEAD_LLM=offline runs the same API with no model: - llm.ts: every LLM call the routes make goes through it and picks its provider per call (llm-mode.ts): gemini.ts by default, offline-llm.ts in offline mode. gemini.ts now creates its client on first use, so importing it needs no key. - offline-llm.ts: the rule-based scorer from packages/scoring for /score and /coach; keyword rules for topics; '' for the LLM-written rewrites, tips, acknowledgements and summaries (the /coach renderers already fall back to their static templates); a template /diff narrative naming the biggest gap; an /improve that asks the static question per weak dimension and appends the answers. POST /onboard/repo/full answers 503 llm_unavailable. - It says so everywhere: GET / reports "llm", /score and /coach return "scorer": "heuristic", coaching text ends with a note, the score card in the extensions shows "Rule-based score", and the API warns at startup. - It never switches on by itself: with no key and no TRAILHEAD_LLM the API still refuses to start (now naming both ways to fix it), so a deploy that lost its key fails loudly instead of quietly scoring with rules. docker-compose.yml no longer requires the key up front, so `TRAILHEAD_LLM=offline docker compose up` works with no .env at all. SELFHOSTING.md documents what offline mode does and does not do, including that the confirming re-score for library promotion is the same rules. Tests: llm-mode and offline-llm unit tests (with fetch stubbed to fail), and an integration test against real Postgres that scores, coaches, promotes, diffs and improves in offline mode and asserts no request-path Gemini call was made (verified: pointing /score back at gemini.ts makes it fail). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…DEMO_TEAM) The seeded demo team's secret (trailhead_demo_acme_2026) is in this repo, and schema.sql always creates the team. On a server reachable from a network, anyone could write to its wiki and spend the operator's Gemini quota through it; the only bound was the per-team and per-IP rate limits, and there was no way to turn it off. TRAILHEAD_DEMO_TEAM=on|off controls it. The default stays on for a local setup, but setting TRAILHEAD_ADMIN_TOKEN (the sign of a networked deploy) turns it off unless TRAILHEAD_DEMO_TEAM=on says otherwise. Off, the demo secret gets 401 with reason demo_team_disabled and a message saying what to do; the team row stays, so turning it back on just works. GET / reports `demo_team`. Integration test: default on; off once an admin token is set (other teams unaffected); an explicit value wins in both directions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…demo up`
The seed wrote a wiki and four library prompts but no prompt activity
("Full §11 seed ... deferred"), and compose never ran it, so a fresh stack
showed an empty skill-arc chart and zeros on the team page.
packages/db/seed.mjs now also writes, for three made-up users, 48 scored
prompts each spread over the last six days with scores that drift upwards
(720 skill_observations, deterministic: fixed-seed PRNG), plus 18 captures
with outcomes. It stays idempotent: the synthetic rows are replaced on every
run, which also moves the window up to "now" (the dashboard shows the last
24 h and the team page the last 7 days). The header says plainly that all of
it is synthetic, and the dashboard's demo-team card no longer calls the
fictional Acme Fintech "the team's actual repo conventions".
docker-compose.yml gains a one-shot `seed` service under the `demo`
profile, started after the API is healthy:
TRAILHEAD_LLM=offline docker compose --profile demo up --build
Verified with that command on a fresh volume: seed exits 0, /team/metrics
reports 720 observations, 3 active users, avg 5.2; the dashboard's
/skill-arc renders 24 hour-buckets and /team shows non-zero cards (headless
Chromium). Running the seed twice leaves the same counts.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI built the API image but never ran the stack, so "docker compose up" could break (schema mount, healthcheck, env wiring) without a red build. The docker job now brings the stack up the way a stranger would, with no key (TRAILHEAD_LLM=offline), runs the demo seed, and checks: GET / is ok and offline, a team registers, a wiki write lands, /score answers with the rule-based scorer, and the seeded demo team has its activity. Logs are printed on failure and the stack is always torn down (down -v). The same commands were run locally against a compose stack before committing (all four jq checks true). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
docker compose up.What changed
Rule-based baseline scorer + eval (
e3c7e2c).packages/scoring/src/heuristic-score.mjsscores the five rubric dimensions with transparent rules over surface features: file paths, identifiers, numbers with units, constraint and output phrasing, and text aimed at the scorer. It is deterministic, has no imports and no Node APIs (so a browser can run it), and is a floor for the model, not a substitute.--scorer heuristic, andeval/holdout.json: 16 prompts with bands, written before the rules were frozen and never tuned on.baseline.test.tsruns the scorer over both sets innpm test.Known misses are listed in
apps/api/eval/README.md. The Gemini scorer has still not been run on either set; that needs a key.Keyless demo mode (
1ea5bd1). WithTRAILHEAD_LLM=offline, routes callllm.ts, which picks per call betweengemini.tsandoffline-llm.ts:/diffand/improveuse templates;503.It says so everywhere:
GET /reportsllm,/scoreand/coachreturnscorer: "heuristic", coaching text ends with a note, the extension score card shows "Rule-based score", and the API warns at startup.It never switches on by itself. Without a key and without
TRAILHEAD_LLM, the API still refuses to start, and the error message names both fixes.gemini.tsnow creates its client lazily, and compose no longer requires the key up front.Demo-team switch (
1e44561).TRAILHEAD_DEMO_TEAM=on|off: default on, but off onceTRAILHEAD_ADMIN_TOKENis set. When off, the demo secret gets401 demo_team_disabled.GET /reports the setting.Seeded demo data (
8297fec).seedservice under--profile demo.CI compose smoke test (
7abf3f2). The docker job runsTRAILHEAD_LLM=offline docker compose up --waitand the seed, then checks withjq: health, a team registration, a wiki write, a rule-based score, and the seeded metrics. It always tears down.Verification
npm run typecheck/npm run lint/npm audit --audit-level=highnpm testllm-mode,offline-llm, ordering,--scorer heuristic, baseline ×3); scoring 43 (+8 heuristic rule tests); score-card 6 (+1)routes/score.tsback atgemini.tsmakes the offline integration test fail.TRAILHEAD_LLM=offline docker compose -p … --profile demo up --build, fresh volumeGET /→llm: offline; seed exits 0./team/metrics: 720 observations, 3 users, avg 5.2./score "fix it"→ 0; a fully specified prompt → 8./coachuses the library example and ends with the offline note.next dev) against that stack, headless Chromium/skill-arcrenders 24 hour-buckets;/teamcards are non-zerojqchecks trueNeeds a human
apps/api/eval/README.md), for exampleGEMINI_API_KEY=… npm --workspace=apps/api run eval -- --yes --runs 5(150 calls). Then compare it with the baseline rows above.TRAILHEAD_PROMOTION_MODE=reviewfor teams that will rely on the library.🤖 Generated with Claude Code