I develop models that connect perception with higher-level understanding. My past work focused on video understanding and visual perception, and I'm now expanding into multimodal models that follow events through time, ground language in visual evidence, and learn coherent reasoning strategies.
I recently completed my Ph.D. in Computer Science at Dartmouth College. I work across video understanding, visual grounding, reinforcement-learning post-training, and transparent reasoning evaluation.
I'm also interested in efficient AI: building smaller, more capable models and the infrastructure and systems that make them practical. This includes high-performance inference systems—continuous batching, multimodal prefix caching, and native model serving—across Apple Silicon, edge devices, and other resource-constrained hardware.
My work has appeared at CVPR, ICCV, ECCV, ICML, WACV, and other leading venues, as well as in PLOS Digital Health, the Journal of Pathology Informatics, Frontiers in Medical Technology, and other journals.
I earned my Ph.D. at Dartmouth College, advised by SouYoung Jin, and founded Wiqonn, an AI lab growing ambitious AI work in Latin America.