
Yong Zheng-XinResearcher
OpenAI
Member of Technical Staff, RSI preparedness/safety
via OpenAI
Researcher

Member of Technical Staff at OpenAI
Korbak studies whether language models reveal enough of their reasoning for people to catch deception or misbehavior. At OpenAI he works on monitoring agents for misalignment, building on earlier research into honest post-training, AI control and using human preferences during pretraining.
See something inaccurate or outdated?
OpenAI
AI Security Institute
AnthropicWork and education history is primarily focused on AI-relevant roles and may not be comprehensive.
Last editorial review: August 9, 2026