跳到正文
稳定技术报道时间 2026-08-25 23:14

Granite 4.2 大型语言模型:增强的架构和训练细节

Granite 已发布其 4.2 大型语言模型的详细信息,包括新的架构、训练方法以及基础设施改进。

01

证据

  • HHugging Face Blog公司披露2026-08-25 23:14
    Overview Model Architecture Pre-Training SFT: Data Preparation & Quality Control Data Quality Control SFT Training Details Phase 2 SFT for the 30B Model Reinforcement Learning: A Multi-Stage, Multi-Environment Pipeline Training Methodology The Staged Curriculum Reward Signals Fo…
    查看来源
  • HHugging Face Blog公司披露2026-08-25 23:14
    TL;DR: Granite 4.2 is our first family of dense, decoder-only reasoning LLMs, released in three sizes: 3B, 8B, and 30B . These models are post-trained from Granite-4.1 base models.
    查看来源
  • HHugging Face Blog公司披露2026-08-25 23:14
    Granite 4.2 models are built on a decoder-only dense transformer architecture with the following core components:
    查看来源
  • HHugging Face Blog公司披露2026-04-29 23:01
    Overview Model Architecture Pre-Training Phase 1: General Pre-Training (10T tokens) Phase 2: Math/Code Pre-Training (2T tokens) Phase 3: High-Quality Data Annealing (2T tokens) Phase 4: High-Quality Data Annealing — Refinement (0.5T tokens) Phase 5: Long Context Training (LCE) S…
    查看来源