Skip to content
View keppy's full-sized avatar
:octocat:
hacking
:octocat:
hacking

Block or report keppy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
keppy/README.md

Keppy's Landing Pad

Hello, I work at Stemuli training math tutoring models.

Research bench where you can find open source software and reports: keppylab.com

Collective research & writing: yokozunasan.com

Recent Projects

titans-mini A streaming engine over a swappable test-time memory core (MLP-weights memory vs. generated-weights vector memory)

thomas thomas.train() — a training harness. Case→reward→train: take a Case set and a score function, get a baseline card (gonogo), run a training loop, compare before/after. Two paths: pretrain (Modal GPU) and post-train (Tinker LoRA RL or TRL GRPO + vLLM).

cotfaith CoT (un)faithfulness, study one: hint-following and confession rates on Qwen3-1.7B — pre-registered decision log, blind-labeled judge validation, byte-exact run artifacts

gonogo The eval harness from my implementations, open sourced. Scores an agent on your real cases and returns a deployment decision, including "not enough evidence yet."

describe• Speak systems into existence

cobol-reporter 🔭 A RAG and report generator for helping to understand COBOL code

Disease Lab 🧪 Knowledge graph informed AI for disease research

WorldEnder.ai 🌎 A text-adventure about extinction events with a RAG backend

Pinned Loading

  1. gonogo gonogo Public

    The eval harness from our implementations, open sourced. Scores an agent on your real cases and returns a deployment decision, including "not enough evidence yet."

    Python

  2. cotfaith cotfaith Public

    CoT (un)faithfulness, study one: hint-following and confession rates on Qwen3-1.7B — pre-registered decision log, blind-labeled judge validation, byte-exact run artifacts

    HTML

  3. recurse recurse Public

    A numbered visual series that remembers itself: plan → draw → render → reflect → remember

    Python

  4. cdiddy77/live-diagrammer cdiddy77/live-diagrammer Public

    TypeScript 1 1