UnimixLM pietrolesci/small_bpe128k Updated Aug 8, 2025 pietrolesci/small_multigram128k Updated Jul 24, 2025 pietrolesci/small_tokmix128k Updated Jul 25, 2025 pietrolesci/small_unigramlm128k Updated Jul 27, 2025
Interesting Pre-Training Datasets Zyphra/Zyda-2 Preview • Updated Aug 6, 2025 • 190k • 95 HuggingFaceTB/dclm-edu Viewer • Updated Mar 7, 2025 • 1B • 8.54k • 32 HuggingFaceFW/fineweb-edu Viewer • Updated Jul 11, 2025 • 3.5B • 635k • 1.1k HuggingFaceTB/stack-edu Viewer • Updated Mar 20, 2025 • 167M • 5.52k • 71
UnimixLM pietrolesci/small_bpe128k Updated Aug 8, 2025 pietrolesci/small_multigram128k Updated Jul 24, 2025 pietrolesci/small_tokmix128k Updated Jul 25, 2025 pietrolesci/small_unigramlm128k Updated Jul 27, 2025
Interesting Pre-Training Datasets Zyphra/Zyda-2 Preview • Updated Aug 6, 2025 • 190k • 95 HuggingFaceTB/dclm-edu Viewer • Updated Mar 7, 2025 • 1B • 8.54k • 32 HuggingFaceFW/fineweb-edu Viewer • Updated Jul 11, 2025 • 3.5B • 635k • 1.1k HuggingFaceTB/stack-edu Viewer • Updated Mar 20, 2025 • 167M • 5.52k • 71