- {SITE.name} is an open community of contributors — engineers, researchers, students,
- and curious learners — writing down what we figure out about modern machine learning systems and
- sharing it with the world.
+ {SITE.name} is an open community of engineers, researchers, students, and curious
+ learners. We share what we learn about machine learning systems so others can build on it.
- The goal is simple: collect honest, technically grounded writing about ML systems — the
- kernels, the schedulers, the embeddings, the deployments — without the marketing gloss.
- Anyone interested in this domain is welcome to write here, whether you've shipped systems
- for years or you just figured something out last week. If you've learned something worth
- passing on, come write with us.
+ You'll find practical explanations and lessons from building ML systems, from kernels and
+ schedulers to training and deployment. We value clear writing, technical depth, and honest
+ accounts of what worked and what didn't.
- We think this knowledge is about to matter much more than it does today —
- here's why this exists.
+ Whether you've built systems for years or just figured something out, your experience can
+ help someone else. Everyone is welcome to contribute.
- You don't need an invitation to contribute. Open the editor, share an
- explanation or experience, and submit it for review. Every article carries its author's
- name; every useful perspective helps the community learn.
+ Open the editor, share an explanation or experience, and submit it for
+ review. Every article is credited to its author.
);
diff --git a/src/content/tools/attention-viz/index.mdx b/src/content/tools/attention-viz/index.mdx
index 3502a8c..2a78243 100644
--- a/src/content/tools/attention-viz/index.mdx
+++ b/src/content/tools/attention-viz/index.mdx
@@ -1,7 +1,7 @@
---
name: Attention Visualizer
-summary: Explore a token-by-token attention map, from causal masking to recurring attention patterns.
-tag: Live
+summary: Explore a synthetic attention map illustrating causal masking and recurring patterns.
+tag: Beta
icon: ◎
authors:
- lchen
@@ -18,23 +18,10 @@ import AttentionViz from './AttentionViz';
-## What it does
+## What it shows
-Loads any model from Hugging Face and renders the attention weights produced for a prompt you control. Built for the moment when you're trying to convince yourself that a specific head is doing what you think it's doing.
-
-## Why it's useful
-
-Reading the attention matrix as raw numbers is hopeless. Reading it as a heatmap with the tokens labeled along both axes makes patterns jump out — causal triangles, sink tokens absorbing residual mass, induction heads aligning along the off-diagonal. A few minutes here often saves hours of grepping through paper figures.
-
-## How to use it
-
-1. Paste a model ID from Hugging Face (e.g. `meta-llama/Llama-3.1-8B`).
-2. Type a prompt — short ones make the visualization legible.
-3. Pick a layer and head from the sidebar.
-4. Watch the matrix render. Hover any cell to see the query/key tokens.
+A synthetic, animated attention matrix for a fixed sentence. Rows represent query positions and columns represent key positions. The triangular shape illustrates causal masking: a token cannot attend to future tokens.
## Limitations
-- Currently runs on a backend GPU, so very large models may queue.
-- Multi-query and grouped-query attention are visualized per query head; KV-shared heads appear repeated.
-- This tool is for understanding, not for high-throughput batch analysis.
+This is a visual demonstration, not measured attention from a trained model. Model loading, custom prompts, layer selection, and head inspection are not implemented. The colors should not be interpreted as evidence about any specific model.
diff --git a/src/content/tools/gpu-mem-calc/GpuMemoryCalc.tsx b/src/content/tools/gpu-mem-calc/GpuMemoryCalc.tsx
index e9023dd..59a4644 100644
--- a/src/content/tools/gpu-mem-calc/GpuMemoryCalc.tsx
+++ b/src/content/tools/gpu-mem-calc/GpuMemoryCalc.tsx
@@ -234,6 +234,7 @@ export default function GpuMemoryCalc({ compact = false }: { compact?: boolean }
>
- Back-of-envelope tokens/sec for a given model, precision, and hardware. Memory-bound
- regime only; assumes batched serving with a healthy KV cache headroom.
+ A weight-bandwidth ceiling, not measured throughput. KV memory uses a fixed example: 80
+ layers, 64 KV heads, head dimension 128, and BF16 cache. Model size changes weights
+ only.
- ⚠ Estimate is memory-bound roofline only. Actual numbers depend on kernel quality,
- continuous batching, speculative decoding, and a dozen other things this tool doesn't
- model.
+ ⚠ The ceiling ignores KV-cache traffic and compute limits. Precision describes weight
+ storage, not native GPU support. Actual numbers depend on kernel quality, continuous
+ batching, speculative decoding, and a dozen other things this tool doesn't model.
diff --git a/src/content/tools/throughput-calc/index.mdx b/src/content/tools/throughput-calc/index.mdx
index fd4949d..26fdf4a 100644
--- a/src/content/tools/throughput-calc/index.mdx
+++ b/src/content/tools/throughput-calc/index.mdx
@@ -1,6 +1,6 @@
---
name: Throughput Calculator
-summary: Estimate memory-bound token throughput and KV-cache size across models, precisions, batch sizes, and GPUs.
+summary: Explore a weight-bandwidth ceiling and a fixed KV-cache memory example across precisions, batch sizes, and GPUs.
tag: Live
icon: ◐
authors:
@@ -18,14 +18,15 @@ import ThroughputCalc from './ThroughputCalc';
## What it does
-Back-of-the-envelope throughput estimation that's accurate enough to inform real architecture decisions. Plug in a model, precision, batch size, and GPU. Get tokens/sec, KV-cache size, and a verdict on whether you'll fit on one card.
+Shows an ideal weight-read ceiling: GPU memory bandwidth divided by weight bytes, multiplied by batch size. It is an upper bound for a simplified decode step, not a throughput prediction.
-## Why it's useful
-
-Most "how fast will this run" questions can be answered without standing up infrastructure. This tool encodes the memory-bound roofline that governs LLM inference, so you can compare options at the cost of one keystroke instead of one cluster-hour.
+The memory example assumes 80 layers, 64 KV heads, a head dimension of 128, and a BF16 cache. These dimensions stay fixed when you change the parameter count. Use the Training Memory Calculator for architecture-specific memory estimates.
## Limitations
-- Memory-bound regime only. Doesn't model the compute-bound prefill ceiling.
-- Doesn't account for continuous batching, speculative decoding, or prefix sharing — those are explicit knobs in real serving stacks.
-- Use the result as a rule of thumb, not a procurement quote.
+- The ceiling omits KV-cache traffic, compute limits, communication, and kernel overhead. Longer sequences increase the displayed memory but do not change this weight-only ceiling.
+- Precision sets weight storage size. It does not establish that a GPU supports native arithmetic at that precision or that suitable kernels exist.
+- Memory fit reserves 15% headroom. A configuration that needs sharding cannot achieve the displayed ceiling on a single GPU; multi-GPU performance is not modeled.
+- Use measured benchmarks for deployment decisions.
+
+See [How To Scale Your Model: Transformer Inference](https://jax-ml.github.io/scaling-book/inference/) for the full bandwidth and compute model.
diff --git a/src/pages/about.astro b/src/pages/about.astro
index 44a5c26..75fd1f6 100644
--- a/src/pages/about.astro
+++ b/src/pages/about.astro
@@ -2,6 +2,7 @@
import BaseLayout from '@/layouts/BaseLayout.astro';
import PageIntro from '@/components/PageIntro.astro';
import FounderBee from '@/components/FounderBee.astro';
+import CommunityIllustration from '@/components/CommunityIllustration.astro';
import IconLinkCard from '@/components/IconLinkCard.astro';
import { SITE } from '@/lib/site';
@@ -36,49 +37,70 @@ const links = [
-
-
- {SITE.name} is an open community of engineers, researchers, students, and curious
- learners. We share what we learn about machine learning systems so others can build on it.
-
+
+
+
+ {SITE.name} is an open community of engineers, researchers, students, and curious
+ learners. We share what we learn about machine learning systems so others can build on it.
+
-
- You'll find practical explanations and lessons from building ML systems, from kernels and
- schedulers to training and deployment. We value clear writing, technical depth, and honest
- accounts of what worked and what didn't.
-
+
+ You'll find practical explanations and lessons from building ML systems, from kernels and
+ schedulers to training and deployment. We value clear writing, technical depth, and honest
+ accounts of what worked and what didn't.
+
-
- Whether you've built systems for years or just figured something out, your experience can
- help someone else. Everyone is welcome to contribute.
-
+
+ Whether you've built systems for years or just figured something out, your experience can
+ help someone else. Everyone is welcome to contribute.
+
-
- Open the editor, share an explanation or experience, and submit it for
- review. Every article is credited to its author.
-
+
+ Open the editor, share an explanation or experience, and submit it
+ for review. Every article is credited to its author.
+
- You can read everything without signing in. The in-browser writing tool at /write runs entirely in your browser — what you type stays on your device, and nothing is sent to us
- until you choose to download it and submit it by pull request, issue, or email.
+ You can read everything without signing in. Drafts in the writing tool at /write
+ are saved on your device. When you submit an article for review, its content, attached files,
+ and author details are sent to our publishing service to create a public pull request on GitHub.
+ You can also download your draft and submit it yourself by pull request, issue, or email.
Analytics
@@ -82,27 +81,16 @@ const canonical = `${SITE.url}/privacy`;
diff --git a/src/pages/tags/[tag]/[...page].astro b/src/pages/tags/[tag]/[...page].astro
index 1097090..66a4515 100644
--- a/src/pages/tags/[tag]/[...page].astro
+++ b/src/pages/tags/[tag]/[...page].astro
@@ -3,6 +3,7 @@ import type { GetStaticPaths, Page } from 'astro';
import { getCollection } from 'astro:content';
import BaseLayout from '@/layouts/BaseLayout.astro';
import Pager from '@/components/Pager.astro';
+import PageIntro from '@/components/PageIntro.astro';
import PostRow from '@/components/PostRow.astro';
import { sortPostsByDate, tagSlug } from '@/lib/data';
import { resolvePostAuthors, type PostWithAuthors } from '@/lib/posts';
@@ -79,16 +80,12 @@ const title =
jsonLd={[collectionJsonLd, breadcrumbJsonLd]}
>
{page.data.map((a) => )}
diff --git a/src/pages/why.astro b/src/pages/why.astro
index bf51875..32927b5 100644
--- a/src/pages/why.astro
+++ b/src/pages/why.astro
@@ -10,7 +10,7 @@ const pageJsonLd = {
'@type': 'WebPage',
name: 'Why this exists',
description:
- 'Why learning machine learning systems matters: intelligence is moving to personal, local devices — and the systems layer decides how fast we get there.',
+ 'Why learning machine learning systems matters: intelligence is moving to personal, local devices, and systems engineering helps make that possible.',
url: canonical,
isPartOf: { '@type': 'WebSite', name: SITE.name, url: SITE.url },
};
@@ -18,76 +18,68 @@ const pageJsonLd = {
-
+
+
+ Understanding the systems behind AI helps more people build, question, and improve them.
+
+
- Every important technology ends up boring. Electricity, databases, GPS — miracles that
- became plumbing. Machine intelligence is on the same path, and we are living through its
- plumbing years. This site is about those years, and the people doing the work.
+ ML systems turn model capabilities into something people can use. This site is for the
+ people learning how that happens and sharing what they discover.
-
Most "model progress" is systems progress
+
Systems make models useful
- Ask what actually changed between the demo that amazed you and the product you use every
- day: tokens got cheaper, first tokens got faster, contexts got longer, models started
- fitting on hardware you own. Almost none of that came from smarter weights. It came from
- quantization, batching, caches, kernels, schedulers — the unglamorous layer underneath.
- The distance between a demo and a product is measured in milliseconds and megabytes, and
- systems engineers are the ones who close it.
+ A capable model is only part of a working product. Memory use, response time, training
+ cost, and reliability matter too. Quantization, batching, caches, kernels, and
+ schedulers help make models practical. We want to make that work easier to understand.
-
Computing always moves closer to you
+
More choice in where AI runs
- Mainframe to desktop, desktop to pocket, cloud to edge — every generation of computing
- ends up nearer to the person using it, because latency, cost, and privacy all pull the
- same way. Intelligence is on the same road. The endpoint is a model that runs on devices
- you own, tuned on your own context — your notes, your work, your family's routines — a
- private intelligence layer that answers to you and no one else. A model that knows you
- that well shouldn't live in someone else's building. The datacenter era of AI is its
- mainframe era, and the people who understand inference at the edge are the ones who will
- end it.
+ Some workloads belong in a datacenter. Others benefit from running on a laptop, phone,
+ or device nearby. Local inference can offer privacy, offline access, and lower latency,
+ with real limits on memory and power. Understanding those tradeoffs gives people more
+ control over the systems they use.
-
The fundamentals outlast the headlines
+
Fundamentals outlast the headlines
- Architectures churn monthly; the systems layer barely moves. Memory hierarchies,
- arithmetic intensity, batching tradeoffs, the cost of moving a byte versus computing on
- it — these were true before transformers and will be true after them. Learning ML
- systems is learning the invariants: knowledge that compounds for decades while the
- leaderboards reshuffle.
+ Models and frameworks change quickly. Memory hierarchies, arithmetic intensity,
+ batching, and the cost of moving data remain useful ways to reason about them. Learning
+ these fundamentals helps you evaluate new ideas instead of starting from scratch each
+ time.
-
The bottleneck is people
+
Knowledge grows when we share it
- The knowledge that makes all of this work is concentrated in a handful of infrastructure
- teams and scattered across conference talks and half-finished blog posts. That scarcity
- is the real constraint on how fast the local, personal future arrives. The fix is old
- and reliable: write things down, in the open, where anyone can learn them. A field grows
- exactly as fast as its commons.
+ Useful knowledge is scattered across papers, code, talks, and individual experience. A
+ clear explanation or an honest account of a failed approach can save someone else days
+ of work. Publishing it openly makes that experience available beyond one team.
-
So we write
+
A place to contribute
- Articles, primers, and tools from practitioners — honest, technically grounded, free to
- read, open to anyone who has figured something out and is willing to pass it on. If the
- future we described sounds right to you, help build the commons that gets us there
- sooner.
+ We bring together articles, primers, and tools from people learning and building ML
+ systems. Everything is free to read, and anyone can submit work for review. If you've
+ learned something worth passing on, there's room for it here.