- Every important technology ends up boring. Electricity, databases, GPS — miracles that
- became plumbing. Machine intelligence is on the same path, and we are living through its
- plumbing years. This site is about those years, and the people doing the work.
+ ML systems turn model capabilities into something people can use. This site is for the
+ people learning how that happens and sharing what they discover.
- Most "model progress" is systems progress
+ Systems make models useful
- Ask what actually changed between the demo that amazed you and the product you use every
- day: tokens got cheaper, first tokens got faster, contexts got longer, models started
- fitting on hardware you own. Almost none of that came from smarter weights. It came from
- quantization, batching, caches, kernels, schedulers — the unglamorous layer underneath.
- The distance between a demo and a product is measured in milliseconds and megabytes, and
- systems engineers are the ones who close it.
+ A capable model is only part of a working product. Memory use, response time, training
+ cost, and reliability matter too. Quantization, batching, caches, kernels, and
+ schedulers help make models practical. We want to make that work easier to understand.
- Computing always moves closer to you
+ More choice in where AI runs
- Mainframe to desktop, desktop to pocket, cloud to edge — every generation of computing
- ends up nearer to the person using it, because latency, cost, and privacy all pull the
- same way. Intelligence is on the same road. The endpoint is a model that runs on devices
- you own, tuned on your own context — your notes, your work, your family's routines — a
- private intelligence layer that answers to you and no one else. A model that knows you
- that well shouldn't live in someone else's building. The datacenter era of AI is its
- mainframe era, and the people who understand inference at the edge are the ones who will
- end it.
+ Some workloads belong in a datacenter. Others benefit from running on a laptop, phone,
+ or device nearby. Local inference can offer privacy, offline access, and lower latency,
+ with real limits on memory and power. Understanding those tradeoffs gives people more
+ control over the systems they use.
- The fundamentals outlast the headlines
+ Fundamentals outlast the headlines
- Architectures churn monthly; the systems layer barely moves. Memory hierarchies,
- arithmetic intensity, batching tradeoffs, the cost of moving a byte versus computing on
- it — these were true before transformers and will be true after them. Learning ML
- systems is learning the invariants: knowledge that compounds for decades while the
- leaderboards reshuffle.
+ Models and frameworks change quickly. Memory hierarchies, arithmetic intensity,
+ batching, and the cost of moving data remain useful ways to reason about them. Learning
+ these fundamentals helps you evaluate new ideas instead of starting from scratch each
+ time.
- The bottleneck is people
+ Knowledge grows when we share it
- The knowledge that makes all of this work is concentrated in a handful of infrastructure
- teams and scattered across conference talks and half-finished blog posts. That scarcity
- is the real constraint on how fast the local, personal future arrives. The fix is old
- and reliable: write things down, in the open, where anyone can learn them. A field grows
- exactly as fast as its commons.
+ Useful knowledge is scattered across papers, code, talks, and individual experience. A
+ clear explanation or an honest account of a failed approach can save someone else days
+ of work. Publishing it openly makes that experience available beyond one team.
- So we write
+ A place to contribute
- Articles, primers, and tools from practitioners — honest, technically grounded, free to
- read, open to anyone who has figured something out and is willing to pass it on. If the
- future we described sounds right to you, help build the commons that gets us there
- sooner.
+ We bring together articles, primers, and tools from people learning and building ML
+ systems. Everything is free to read, and anyone can submit work for review. If you've
+ learned something worth passing on, there's room for it here.
@@ -194,14 +186,14 @@ const pageJsonLd = {
}
.why-lede {
font-family: var(--font-read);
- font-style: italic;
+ font-style: normal;
font-size: 21px;
line-height: 1.6;
color: var(--ink);
margin: 8px 0 0;
}
.why-body section {
- margin-top: 48px;
+ margin-top: 36px;
}
.why-body h2 {
font-family: var(--font-display);
@@ -221,6 +213,6 @@ const pageJsonLd = {
display: flex;
gap: 12px;
flex-wrap: wrap;
- margin-top: 56px;
+ margin-top: 36px;
}
diff --git a/src/styles/publication.css b/src/styles/publication.css
index 24deafd..83bb3f3 100644
--- a/src/styles/publication.css
+++ b/src/styles/publication.css
@@ -131,6 +131,9 @@ main {
padding-top: 56px;
}
.footer-tagline {
+ font-family: var(--font-sans);
+ font-style: normal;
+ font-weight: 400;
max-width: 320px;
font-size: 15px;
line-height: 1.7;
@@ -597,32 +600,66 @@ input[type='radio'] {
}
.publication-topic-chips {
display: flex;
- gap: 10px;
- overflow-x: auto;
- padding: 3px 3px 6px;
- scrollbar-width: thin;
- scrollbar-color: var(--line-2) transparent;
+ flex-wrap: wrap;
+ align-items: baseline;
+ font-family: var(--font-read);
+ font-size: 17px;
+ letter-spacing: -0.01em;
+ color: var(--ink-2);
}
.publication-topic-chips a {
- flex: 0 0 auto;
- padding: 9px 16px;
- border: 1px solid var(--line);
- border-radius: 999px;
- color: var(--ink-2);
- font-size: 14px;
- line-height: 1.4;
+ position: relative;
+ color: inherit;
+ padding: 2px 0;
text-decoration: none;
- transition:
- border-color 150ms,
- color 150ms;
+ transition: color 150ms;
+}
+.publication-topic-chips a::before {
+ content: '';
+ position: absolute;
+ inset: -4px -7px;
+ border: 1px solid color-mix(in srgb, var(--accent) 35%, var(--line));
+ border-radius: 6px;
+ opacity: 0;
+ pointer-events: none;
+ transition: opacity 160ms ease;
+}
+.publication-topic-chips a:not(:last-child) {
+ margin-right: 32px;
+}
+.publication-topic-chips a:not(:last-child)::after {
+ content: '·';
+ position: absolute;
+ right: -18px;
+ color: var(--ink-4);
+}
+.publication-topic-chips a:hover::before,
+.publication-topic-chips a:focus-visible::before {
+ opacity: 1;
}
.publication-topic-chips a:hover {
- border-color: var(--accent);
color: var(--accent);
}
.publication-topic-chips a:focus-visible {
- outline: 2px solid var(--accent);
- outline-offset: 1px;
+ outline: none;
+ color: var(--accent);
+}
+.publication-topic-chips a:focus-visible::before {
+ border-color: var(--accent);
+}
+@media (prefers-reduced-motion: reduce) {
+ .publication-topic-chips a,
+ .publication-topic-chips a::before {
+ transition: none;
+ }
+}
+.publication-chip-tag {
+ margin-left: 8px;
+ font: 10px/1 var(--font-mono);
+ letter-spacing: 0.08em;
+ text-transform: uppercase;
+ color: var(--ink-3);
+ vertical-align: middle;
}
.nav-topics {
@@ -866,3 +903,12 @@ input[type='radio'] {
transition: none;
}
}
+
+/* Keep schematic frames legible against the dark hero without changing data colors. */
+[data-theme='dark'] .hero-scenes {
+ --line: color-mix(in srgb, var(--ink-3) 48%, var(--paper));
+ --line-2: color-mix(in srgb, var(--ink-3) 68%, var(--paper));
+}
+[data-theme='dark'] .hero-scenes line[stroke='var(--ink-3)'][stroke-width='0.5'] {
+ stroke-width: 0.85;
+}