Granite 4.2 LLMs: Enhanced Architecture and Training Details
Granite has released details on their 4.2 Large Language Models, including a new architecture, training methodology, and infrastructure improvements.
Granite has released details on their 4.2 Large Language Models, including a new architecture, training methodology, and infrastructure improvements.
Overview Model Architecture Pre-Training SFT: Data Preparation & Quality Control Data Quality Control SFT Training Details Phase 2 SFT for the 30B Model Reinforcement Learning: A Multi-Stage, Multi-Environment Pipeline Training Methodology The Staged Curriculum Reward Signals Fo…
TL;DR: Granite 4.2 is our first family of dense, decoder-only reasoning LLMs, released in three sizes: 3B, 8B, and 30B . These models are post-trained from Granite-4.1 base models.
Granite 4.2 models are built on a decoder-only dense transformer architecture with the following core components:
Overview Model Architecture Pre-Training Phase 1: General Pre-Training (10T tokens) Phase 2: Math/Code Pre-Training (2T tokens) Phase 3: High-Quality Data Annealing (2T tokens) Phase 4: High-Quality Data Annealing — Refinement (0.5T tokens) Phase 5: Long Context Training (LCE) S…