Llama 2
model Your tags
Your notes
7B, 13B, and 70B parameter dense Transformers trained on 2T tokens. 70B uses Grouped Query Attention (GQA). 4K token context. First Llama with commercial licensing. Chat variants fine-tuned with RLHF.
Llama 2 was the first open-weight model with a permissive commercial license, enabling enterprise adoption at scale. The 70B chat model was competitive with early GPT-3.5. AA Intelligence Index: 3 (70B). By Touvron et al.
Model Details
Architecture DENSE
Parameters 70B
Context window 4,096
Training tokens 2T
AA Intelligence 3
Variants
| Name | Parameters | Notes |
|---|---|---|
| Llama 2 7B | 7B | — |
| Llama 2 13B | 13B | — |
| Llama 2 70B | 70B | — |