Companion artifact for Starfish: Fault-Tolerant Far Memory with Low Resource and Performance Overhead (ATC 2026; paper supplied with the AE submission). By Yanwen Xia, Benyong Deng, Hong Huang, Yuzheng Wang, Quanxi Li, Xingda Wei, Xiaobing Feng, Huimin Cui, and Chenxi Wang.
Our code uses Apache-2.0; third-party code and inputs retain their own licenses, including GPL-3.0 for libfibre.
Reproducing this system requires multiple machines connected by 100 Gbps InfiniBand (up to nine machines in the paper, including recovery standbys).
The Starfish runtime, the Non-FT, Hydra-like and Carbink-like baselines, and the evaluation environment are ready. See Quick Start for setup and a minimal experiment.
AE reviewers should provide an SSH public key. Server access for AE reviewers will be provided through WireGuard.
Paths below are relative to artifact-evaluation/:
| Component | Paper relationship |
|---|---|
runtime/starfish/, runtime/nonft/, runtime/hydra/, runtime/carbink/ |
Starfish, Non-FT, and our Hydra-like and Carbink-like implementations |
apps/, configs/ |
Evaluation applications and configurations |
scripts/figure9/ |
Application-performance experiments, result collection and Figure 9 |
scripts/figure10/–figure13/, scripts/appendix/ |
Latency, resource cost, compute overhead, recovery and appendix figures |
third_party/libfibre/ |
Pinned libfibre source and its errnoname dependency; upstream licenses and source information are retained |
| Guide | Contents |
|---|---|
| Installation | Dependencies, installation and build commands |
| Data and models | Download sources, preparation and input paths |
| Multi-server configuration | Compute/memory roles, SSH/TCP addresses and ports |
| Server table | Server IPs, IB devices and NUMA placement |
| Experiments | Experiment commands, result files and plotting |
| Script index | Figure-specific instructions and CSV formats |
From the repository root:
cd artifact-evaluation
bash scripts/common/setup_environment.sh
bash scripts/common/check_environment.sh
bash scripts/common/build.shUse the prepared data provided under /data/starfish-ae/:
bash scripts/common/use_prepared_data.shThe script records input paths in data/site.json. Set the server addresses
once using multi-server configuration.
To download and prepare inputs from their original sources instead, follow
Data and models.
Run LLaMA at 25% local memory once each with Non-FT, Starfish, and Starfish with recovery from one memory-service failure:
bash scripts/run_fast_check.shThe script builds the required targets and checks that all three runs finish with the same chat output. See fast-check details for configuration and result files.
Start Figure 9 with:
bash scripts/figure9/run.shThe minimal working example uses one compute server and one memory server.
Hardware reference and operating-system requirements:
| Role | CPU / RAM | OS requirement |
|---|---|---|
| Compute | 2× Xeon Gold 6342, 48 cores / 256 GiB | Ubuntu 22.04 |
| Memory | 2× Xeon Silver 4316, 40 cores / 256 GiB | Ubuntu 22.04 |
No specific Linux kernel version is required; compatible RDMA drivers are needed.
Compute tools: GCC 13.1, CMake 3.22.1, ConnectX-5 Ex and RDMA userspace
2410mlnx54-1.2410068. Build the memory server locally if its libc differs.
Dependencies and pinned libfibre/HdrHistogram builds are in the
installation guide. RDMA/SSH access
and 2 MiB HugePages must be configured before running.
| Application | Dataset / model |
|---|---|
| LLaMA | LLaMA 2 7B Chat (FP32) |
| BFS | Friendster social graph |
| MG | NAS Parallel Benchmarks, Class D |
| WordCount | English Wikipedia |
| KV-B | YCSB Workload B: 95% reads, 5% updates; 1 billion operations |
| KV-A | YCSB Workload A: 50% reads, 50% updates; 1 billion operations |
| KV-S | Synthetic: 5% reads, 95% updates; 1 billion operations |
| NQ | Friendster social graph; 2-hop neighborhood queries, 2 million queries |
KV uses 512-byte records and Zipfian skew 0.99. Download sources and preparation instructions are in Data and models.
KV records and requests, NQ queries, and MG grids are generated at runtime using the workload parameters specified in the paper; NQ uses the prepared Friendster graph.
Paper-scale runs use 24 application cores and 256 GB RAM per node with 100 Gbps InfiniBand: two nodes for the Non-FT example, seven for steady-state RS(4,2) experiments, and up to nine including recovery standbys. Reserve about 200 GiB of compute-side workspace for inputs, builds and logs.
The following estimates cover serial application work for three repetitions. Installation, input loading, warm-up and service resets are additional.
| Experiment | Estimated work time |
|---|---|
| Application performance (Fig. 9) | 11.8 hours |
| Tail latency (Fig. 10) | 5.1 hours |
| FT resource cost (Fig. 11) | 3.2 hours |
| Compute-node overhead (Fig. 12) | 1.6 hours |
| Failure recovery (Fig. 13) | 10 minutes |
Please allow 2–3 days for Figure 9.
From artifact-evaluation/:
| Figure | Run / collect | Plot |
|---|---|---|
| Figure 9: application performance | bash scripts/figure9/run.sh |
bash scripts/plot.sh figure9 |
| Figure 10: tail latency | bash scripts/figure10/run.sh |
bash scripts/plot.sh figure10 --input data/figure10.csv |
| Figure 11: FT resource cost | bash scripts/figure11/collect.sh --logs-root results/figure9 |
bash scripts/plot.sh figure11 --input data/figure11.csv |
| Figure 12: compute-node overhead | bash scripts/figure12/collect.sh --logs-root results/figure9 |
bash scripts/plot.sh figure12 --input data/figure12.csv |
| Figure 13: failure recovery | bash scripts/figure13/collect.sh --logs-root results/figure13 |
bash scripts/plot.sh figure13 --logs-root results/figure13 |
The Figure 9 runner writes data/figure9.csv. Input formats and output files
are described in the linked figure guides.
Runs start and stop memory services: use unused ports and dedicated result directories on authorized hosts. Failure injection must target isolated services, never production endpoints.
See the scripts guide for detailed server setup, experiment options, and CSV formats.