"Seq vs Seq: An Open Suite of Paired Encoders and Decoders": the first collection of encoder-only and decoder-only models trained from scratch with identical data, architecture, and recipe — six sizes from 17M to 1B, each in both flavors, on 2T tokens of open data (DCLM, Dolma v1.7, code, scientific papers). A clean testbed for the encoder-vs-decoder question: encoders win MNLI even against larger decoders, decoders win generation, and cross-objective continued training doesn't close the gap. JHU-CLSP (Weller et al.), MIT-licensed with released data and checkpoints.

Model Details

Architecture DENSE
Parameters 1B
Training tokens 2T
License MIT

Variants

Name Parameters Notes
Ettin 17M 17M —
Ettin 32M 32M —
Ettin 68M 68M —
Ettin 150M 150M —
Ettin 400M 400M —
Ettin 1B 1B —

Paper

open-weightpretrainingresearch

Related