A study of sequence weighting at scale
Summary
Jane Street's study investigates data weighting in training LLMs and finds that the effective sequence weight exponent scales non-monotonically across model sizes. It reports aberrant scaling—small models learn general patterns independent of weight, medium models learn patterns proportional to data weight, and large models again learn all patterns regardless of weight—along with potential remedies and implications for scaling experiments.