Drop-Upcycling
updated
9B
•
Updated
•
7
9B
•
Updated
•
6
19B
•
Updated
•
9
9B
•
Updated
•
7
9B
•
Updated
•
7
0.4B
•
Updated
•
9
0.4B
•
Updated
•
10
0.4B
•
Updated
•
8
0.4B
•
Updated
•
18
9B
•
Updated
•
5
19B
•
Updated
•
9
0.4B
•
Updated
•
11
9B
•
Updated
•
6
0.4B
•
Updated
•
16
•
1
2B
•
Updated
•
6
0.2B
•
Updated
•
17
4B
•
Updated
•
6
14B
•
Updated
•
6
llm-jp/Dense-btx-code-expert-152M
0.2B
•
Updated
•
6
•
1
llm-jp/Dense-btx-english-expert-1.5B
2B
•
Updated
•
4
llm-jp/Dense-btx-code-expert-1.5B
2B
•
Updated
•
4
•
1
llm-jp/Dense-btx-japanese-expert-1.5B
2B
•
Updated
•
8
•
1
llm-jp/Dense-btx-english-expert-152M
0.2B
•
Updated
•
6
llm-jp/Dense-btx-japanese-expert-152M
0.2B
•
Updated
•
5
Drop-Upcycling: Training Sparse Mixture of Experts with Partial
Re-initialization
Paper
•
2502.19261
•
Published
•
6
Text Generation
•
73B
•
Updated
•
234
llm-jp/llm-jp-3-8x13b-instruct3
Text Generation
•
73B
•
Updated
•
70
•
8
llm-jp/llm-jp-3-8x1.8b-instruct3
Text Generation
•
9B
•
Updated
•
153
•
4
Text Generation
•
9B
•
Updated
•
14
llm-jp/llm-jp-3-8x13b-instruct2
Text Generation
•
73B
•
Updated
•
25
llm-jp/llm-jp-3-8x1.8b-instruct2
Text Generation
•
9B
•
Updated
•
37
llm-jp/llm-jp-3.1-8x13b-instruct4
Text Generation
•
73B
•
Updated
•
484
•
4
Text Generation
•
73B
•
Updated
•
367