
Peiyi WangResearcher
DeepSeek
Researcher
via DeepSeek
Researcher
Research Scientist at DeepSeek
Shao is a researcher at DeepSeek known for DeepSeekMath and for co-developing Group Relative Policy Optimization (GRPO), a reinforcement-learning method central to recent reasoning models. He is a research scientist at DeepSeek.
See something inaccurate or outdated?
DeepSeekWork and education history is primarily focused on AI-relevant roles and may not be comprehensive.
Last editorial review: August 13, 2026