(无标题)
K2.5 went through a long post-training process to really unleash the potential of the base model.
Using SFT on text alone to bootstrap vision RL, and seeing vision RL improve text performance, made me rethink how generalization really works.
K2.5 went through a long post-training process to really unleash the potential of the base model.
Using SFT on text alone to bootstrap vision RL, and seeing vision RL improve text performance, made me rethink how generalization really works.