
Juan Felipe Cerón UribeResearcher
OpenAI
AI Alignment Research Engineer
via OpenAI
Researcher

Member of Technical Staff at OpenAI
Wallace co-created universal adversarial triggers: short phrases that reveal systematic weaknesses in language models. At OpenAI he helped develop instruction hierarchy, which teaches models to follow trusted system instructions over malicious prompts, and works on privacy, memorisation, and alignment training.
See something inaccurate or outdated?
OpenAI
Google DeepMind
MetaWork and education history is primarily focused on AI-relevant roles and may not be comprehensive.
Last editorial review: August 8, 2026