Paper Reading and Result Reproduction
Prerequisites: lesson 5 experimental method and one implemented systems topic. Budget: 40-60 hours. Outcome: reproduce one bounded claim and separate the paper's evidence from your own.
Diagnostic
Take a benchmark chart and state its workload, baseline, measured quantity, and an unsupported conclusion someone might draw from it. If these are missing, complete the performance studio before reading a large stack of papers.
Read to recover an argument
First identify the problem, claimed contribution, and assumptions. Then inspect the mechanism and evaluation: what changes, what stays fixed, and which comparison supports the conclusion? Finally reconstruct a central result or argument without the paper open.
Record the title, authors, publication venue/year, precise version or URL, and location of the selected claim. A local filename is a discovery aid, not sufficient provenance. Verify it against the title page and an author or publisher source before citing it. Notes in Resources are useful navigation but are not substitutes for the original paper's evidence.
Worked example: narrow the claim before reproducing it
A paper reports that partitioning a workload reduces elapsed time on its cluster. A laptop reproduction with four local workers can investigate the scheduling mechanism and correctness, but it does not reproduce the cluster's absolute speedup.
State a narrower hypothesis: “For this fixed dataset and mapper, partitioning across two local workers changes elapsed time relative to one worker.” Keep parsing, input data, output checks, and measurement boundaries equal. Test a small workload where process startup may dominate and a larger one where parallel work may help. A negative result can be informative if the conditions and differences are explained.
For a MapReduce-style word count, map emits (word,1), shuffle groups by word, and reduce sums counts. The sequential oracle uses the same tokenization. A partition with no words, repeated words across partitions, and a retry of a completed task expose correctness and deduplication assumptions before performance is measured.
Guided assignment
Choose one paper matched to your artifact:
- MapReduce for a small dataflow executor; locate the original through Google Research.
- Raft for a bounded state-machine or election experiment; use the paper's project site.
- A local compiler, memory-safety, or static-analysis paper, after verifying its identity and accessible original source.
Write a one-page preregistered plan containing the selected claim, baseline, variables, seeds, trial protocol, correctness oracle, success/failure interpretation, and available compute limit. “Preregistered” here means committed before running the experiment, not registered with an external research organization.
Reproduce one table, mechanism, counterexample, or figure at an explicitly smaller scale. Then change one variable and report both results. Keep failed attempts, dependency versions, commands, raw measurements, and analysis code. If original code or data is unavailable, label your work a reimplementation and explain what comparison remains possible.
Acceptance: another person can rerun the experiment; output correctness is independently checked; the conclusion is no broader than the observations; differences from the original setup are explicit. A summary of the paper alone does not pass.
Independent transfer and defense
Your reproduction contradicts the paper. Give three hypotheses before declaring the paper wrong. Check: implementation mismatch, workload/environment difference, and statistical or measurement variation are candidates to investigate, not automatic excuses. Design a discriminating next experiment.
Use the local How to Read a Paper as a process aid after checking its title-page metadata. Keep the resulting critique respectful and specific: claim, evidence, limitation, next experiment. Review with the assessment contract.