Quick answer: Step3-VL-10B is StepFun's 10B-parameter vision-language model, scoring 84.0% on MathVista's visual math reasoning benchmark.
Where Step3-VL-10B leads
Where it lags